TLDRocket
Sign in

OpenAI and Anthropic identify significant flaws in SWE-Bench Pro coding evaluation benchmark

Benchmark result Provisional 72% confidence first seen

OpenAI and Anthropic independently audited SWE-Bench Pro, a widely-used benchmark for evaluating coding AI systems, and found substantial issues including ambiguous test cases, inconsistent evaluation criteria, and approximately 30% of tasks being broken due to overly strict tests and unclear specifications. The findings prompted both companies to question the benchmark's reliability and led to calls for the community to reconsider which benchmarks should be trusted for measuring software engineering AI capabilities.

Source coverage

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.