SWE-Bench Pro
This profile is built automatically from TLDRocket coverage.
Specifications
No specifications recorded yet.
Latest developments
You only need the frontier model for one single edit
TLDR Dev · 1 week ago ·
5
The Sequence AI of the Week #895: OpenAI's Show Us Where Coding Evals Break
TheSequence · 2 weeks ago ·
4
Separating signal from noise in coding evaluations
TLDR Dev · 3 weeks ago ·
3
Separating signal from noise in coding evaluations
OpenAI Blog · 3 weeks ago ·
18
Why we no longer evaluate SWE-bench Verified
OpenAI Blog · 5 months ago ·
26
2026
OpenAI and Anthropic identify significant flaws in SWE-Bench Pro coding evaluation benchmark Benchmark result