Why we no longer evaluate SWE-bench Verified
OpenAI Blog 4 months ago
SWE-bench Verified, a benchmark used to evaluate AI coding abilities, has become unreliable due to contamination and flawed test design that misrepresents actual progress. The benchmark's tests have leaked into training data and contain methodological problems that produce inaccurate measurements of frontier model performance. Researchers are now recommending SWE-bench Pro as an alternative evaluation method instead.