SWE-Bench Pro
Model ● Covered in 5 stories + Follow
SWE-Bench Pro is a coding evaluation benchmark designed to assess the performance of AI models on software engineering tasks. Recent audits by OpenAI and Anthropic identified that approximately 30% of its 731 public tasks contain defects such as overly strict tests, unclear specifications, and inconsistent evaluation criteria, leading both organizations to withdraw their recommendations for its use as a standard evaluation tool. The benchmark's validity issues have prompted the AI community to reconsider which coding benchmarks provide reliable measurements of software engineering AI capabilities.
Updated 5 August 2026
Specifications
No specifications recorded yet.
Latest developments
Separating signal from noise in coding evaluations
OpenAI · 2 months ago ·
6
Separating signal from noise in coding evaluations
OpenAI · 2 months ago ·
19
Why we no longer evaluate SWE-bench Verified
OpenAI · 6 months ago ·
29
July 2026
OpenAI and Anthropic identify significant flaws in SWE-Bench Pro coding evaluation benchmark Benchmark result
- You only need the frontier model for one single edit
- The Sequence AI of the Week #895: OpenAI's Show Us Where Coding Evals Break
- Separating signal from noise in coding evaluations
- Separating signal from noise in coding evaluations