TLDRocket
Sign in

Why we no longer evaluate SWE-bench Verified

OpenAI

OpenAI just quietly dropped SWE-bench Verified, the coding benchmark everyone's been quoting for a year. Turns out the test got contaminated and stopped measuring real progress.

Based on reporting by OpenAI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

For a while, SWE-bench Verified was the number people threw around whenever they wanted to prove a model could actually code. OpenAI's own post-mortem now says that number stopped meaning much. The benchmark, built from real GitHub issues, has apparently been leaking into training data long enough that models may simply be pattern-matching solutions they've seen before rather than reasoning through fresh bugs.

That's not the only problem OpenAI flagged. Their analysis found flawed test cases baked into the benchmark itself, meaning some models could pass not because they solved the underlying issue correctly but because the grading criteria were loose or outright wrong. When the yardstick has cracks in it, every score built on top of it inherits the distortion. And once a benchmark gets famous enough to matter for bragging rights, the incentive to memorize it rather than generalize past it only grows.

OpenAI's fix is to stop reporting SWE-bench Verified results going forward and point people toward SWE-bench Pro instead, a version presumably built with more care around contamination and test validity. It's a quiet admission dressed up as a technical note, but the implication is bigger than one leaderboard: a lot of the coding progress claims from the last year rested partly on a test that had already started to rot.

The timing also says something about where the field is. Benchmarks used to have a long shelf life because models weren't good enough to fully saturate them. Now labs are burning through evaluation suites fast enough that contamination and staleness are becoming a routine maintenance problem, not a rare embarrassment.

My take — AI-written commentary, not fact-checked reporting

I've said for a while that benchmark contamination is the industry's open secret nobody wants to date first, so credit to OpenAI for actually naming it instead of quietly padding another SWE-bench number into a launch post. But let's be honest about the pattern: benchmarks get gamed, labs eventually admit it once the number stops being useful for marketing, then everyone migrates to the next shiny eval and the cycle resets. Until independent, closed-off evaluation becomes the norm rather than self-reported honor-system testing, I'd treat every 'state of the art' coding claim as a rough vibe check, not a scientific result.

Read more about this at: OpenAI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.