TLDRocket
Sign in

Large Language Models

268 summarised stories about Large Language Models, each linking back to the original source. Browse all topics →

Monday, 23 February 2026

Why we no longer evaluate SWE-bench Verified

OpenAI Blog 4 months ago

SWE-bench Verified, a benchmark used to evaluate AI coding abilities, has become unreliable due to contamination and flawed test design that misrepresents actual progress. The benchmark's tests have leaked into training data and contain methodological problems that produce inaccurate measurements of frontier model performance. Researchers are now recommending SWE-bench Pro as an alternative evaluation method instead.