Introducing SWE-bench Verified
OpenAI Blog
SWE-bench Verified is a human-validated subset of the SWE-bench dataset designed to more reliably measure how well AI models can solve real-world software engineering problems. The dataset filters the original SWE-bench to remove issues with unclear specifications or incorrect solutions, creating a more trustworthy evaluation benchmark. This enables more accurate comparison of AI coding assistants' performance on genuine bug-fixing and feature implementation tasks.
Why it matters
We’re releasing a human-validated subset of SWE-bench that more reliably evaluates AI models’ ability to solve real-world software issues.