Why we no longer evaluate SWE-bench Verified
OpenAI Blog
SWE-bench Verified, a benchmark used to evaluate AI coding abilities, has become unreliable due to contamination and flawed test design that misrepresents actual progress. The benchmark's tests have leaked into training data and contain methodological problems that produce inaccurate measurements of frontier model performance. Researchers are now recommending SWE-bench Pro as an alternative evaluation method instead.
Why it matters
SWE-bench Verified is increasingly contaminated and mismeasures frontier coding progress. Our analysis shows flawed tests and training leakage. We recommend SWE-bench Pro.