DeepSWE
Reported model scores on DeepSWE, best parseable score first. Each row keeps its verbatim score, test conditions and provenance, and links to the source coverage it was extracted from. Benchmark profile →
| Model | Score | Conditions | Provenance | Measured | Source |
|---|---|---|---|---|---|
| DeepSeek V4 Pro 0813 | 88.5% | pass@4 | independent | — | coverage → |
| GLM-5.3 + GPT-5.6 Sol cascade (routing GLM-5.3 -> Sol when tests fail) | 85.9% | — | vendor-reported | — | coverage → |
| GPT-5.6 Sol | 85.8% | pass@4 | independent | — | coverage → |
| DeepSeek V4 Pro 0813 + Claude Fable 5 | 82.7% | cascading strategy | independent | — | coverage → |
| GLM-5.3 | 81.1% | pass@2 | vendor-reported | — | coverage → |
| Claude Fable 5 | 77.1% | pass@2 | vendor-reported | — | coverage → |
| GPT-5.6 Sol | 72.7% | — | vendor-reported | — | coverage → |
| GPT-5.6 Sol | 72.7% | pass@1 | independent | — | coverage → |
| Claude Fable 5 | 69.7% | pass@1 | vendor-reported | — | coverage → |
| Claude Fable 5 | 69.7% | pass@1 | independent | — | coverage → |
| GLM-5.3 | 69.0% | pass@1 | vendor-reported | — | coverage → |
| DeepSeek V4 Pro 0813 | 62.8% | pass@1 | independent | — | coverage → |
Scores are only comparable within one benchmark under matching conditions — results under different test setups, and scores from other benchmarks, are not directly comparable.