Introducing SWE-bench Verified
OpenAI 2 years ago 50
SWE-bench Verified is a human-validated subset of the SWE-bench dataset designed to more reliably measure how well AI models can solve real-world software engineering problems. The dataset filters the original SWE-bench to remove issues with unclear specifications or incorrect solutions, creating a more trustworthy evaluation benchmark. This enables more accurate comparison of AI coding assistants' performance on genuine bug-fixing and feature implementation tasks.