TLDRocket
Sign in

DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

Together AI Covered by 2 sources

DeepSeek V4 Pro 0813 and Claude Fable 5 were tested on DeepSWE, a software engineering benchmark, revealing a 90x cost difference ($0.24 vs $21.63 per rollout) with Fable leading 69.7% to 62.8% on first attempt but Pro matching or exceeding Fable at higher retry counts. A cascading strategy—running Pro first and escalating to Fable only on failures—achieves 82.7% accuracy at $8.28 per task, outperforming either model alone by 13 percentage points while costing 62% less than Fable standalone. The models disagree on 0.39 correlation, making them complementary: Fable excels in Rust and data serialization while Pro handles concurrency and stateful reactivity better.

Why it matters

We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.