OpenAI and Anthropic identify significant flaws in SWE-Bench Pro coding evaluation benchmark
Benchmark result Provisional 72% confidence first seen
OpenAI and Anthropic independently audited SWE-Bench Pro, a widely-used benchmark for evaluating coding AI systems, and found substantial issues including ambiguous test cases, inconsistent evaluation criteria, and approximately 30% of tasks being broken due to overly strict tests and unclear specifications. The findings prompted both companies to question the benchmark's reliability and led to calls for the community to reconsider which benchmarks should be trusted for measuring software engineering AI capabilities.