SWE-bench Verified
Model ● Covered in 4 stories + Follow
SWE-bench Verified is a human-validated benchmark dataset designed to measure AI models' ability to solve real-world software engineering tasks by filtering the original SWE-bench to remove unreliable test cases. The benchmark has been widely used to evaluate coding agents, with models like a fine-tuned Qwen 32B achieving 59.4% pass@1, though recent analysis indicates the benchmark has become contaminated with training data leakage and methodological flaws that undermine its reliability as an evaluation metric.
Updated 8 August 2026
Specifications
No specifications recorded yet.
Latest developments
Q3 2026
Development of autonomous agent harness frameworks for improving large language model reliability and performance Feature update
Q1 2026
- CoderForge-Preview: SOTA open dataset for training efficient coding agents
- Why we no longer evaluate SWE-bench Verified