TLDRocket
Sign in

Separating signal from noise in coding evaluations

OpenAI Covered by 2 sources

Anthropic checked SWE-Bench Pro, a popular coding benchmark, and found roughly 30% of its tasks are just broken. They've pulled their earlier advice to switch over to it.

Based on reporting by OpenAI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Benchmarks are supposed to be the boring, trustworthy part of AI development — the ruler you use to measure whether a model actually got better at writing code. So it's a little unsettling that Anthropic just audited SWE-Bench Pro and found that nearly a third of its tasks don't hold up under scrutiny.

The problems aren't exotic. Anthropic's team flagged overly strict test cases that fail correct solutions for trivial reasons, prompts that leave out details a human engineer would need to actually solve the problem, and instructions that point coding agents in the wrong direction entirely. In other words, roughly 30% of the time, a model could produce genuinely good code and still get marked wrong — or produce mediocre code and get lucky because the test was too loose to catch it.

This matters because SWE-Bench Pro had been gaining traction as the harder, more rigorous successor to the original SWE-Bench, which itself became somewhat compromised once labs started training against it. Anthropic had reportedly recommended teams move to the Pro version as a cleaner signal. Now that recommendation is retracted, which is a notable reversal — it's not every day a major lab publicly walks back guidance about which benchmark to trust.

The deeper issue here isn't really about one dataset. It's a reminder that as coding agents get evaluated on increasingly complex, real-world-style repositories, the benchmarks themselves need the same engineering rigor as the models being tested. A single mislabeled edge case in a multiple-choice quiz is a rounding error. A single broken integration test in an agentic coding benchmark can silently cap a model's measured performance, or inflate it, in ways that ripple into leaderboards, marketing claims, and research decisions.

Anthropic says it's sharing the audit findings so the benchmark's maintainers can fix the flagged tasks, which is the right instinct. But it also means anyone citing SWE-Bench Pro scores over the past few months should probably squint at those numbers a little harder.

My take — AI-written commentary, not fact-checked reporting

This is exactly the kind of housekeeping the field doesn't do often enough — labs love publishing benchmark wins, way less eager to audit the benchmarks themselves. I'd rather see three solid, verified evals than ten flashy ones nobody's stress-tested, and I hope this pushes other benchmark maintainers to invite the same scrutiny before their numbers end up in someone's earnings call.

Read more about this at: OpenAI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.