TLDRocket
Sign in

Why we no longer evaluate SWE-bench Verified

OpenAI Blog

SWE-bench Verified, a benchmark used to evaluate AI coding abilities, has become unreliable due to contamination and flawed test design that misrepresents actual progress. The benchmark's tests have leaked into training data and contain methodological problems that produce inaccurate measurements of frontier model performance. Researchers are now recommending SWE-bench Pro as an alternative evaluation method instead.

Why it matters

SWE-bench Verified is increasingly contaminated and mismeasures frontier coding progress. Our analysis shows flawed tests and training leakage. We recommend SWE-bench Pro.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.