DeepSeek V4 Pro benchmarked against Claude Fable 5 and GPT-5.6 Sol on DeepSWE software engineering tasks
Benchmark result Provisional 92% confidence first seen
Researchers from Together AI evaluated DeepSeek V4 Pro 0813 against two proprietary AI models (Claude Fable 5 and GPT-5.6 Sol) on the DeepSWE benchmark, a software engineering task suite with 113 problems. The analysis demonstrated that DeepSeek Pro, while less accurate at single-attempt success, significantly undercuts competitors on cost (90x-35x cheaper) and can match or exceed their performance when given multiple retry attempts. A cascading inference strategy that runs the cheaper model first and escalates to premium models only on failure achieved superior cost-performance tradeoffs compared to using either model alone.