TLDRocket
Sign in

Kimi K3 vs Claude Fable 5 on DeepSWE: Cost and Coding

Together AI Covered by 33 sources

Kimi K3, a new open-weight coding model, nearly ties Claude Fable 5 on a tough software-engineering benchmark. It costs a third as much and gets even better if you let it try more than once.

Benchmarks usually crown a single winner and everyone moves on. This one is more interesting because the runner-up is the story. Kimi K3, the new open-weight model from Moonshot AI, landed on DeepSWE on July 16 with 452 graded rollouts, and it lands just 1.4 points behind Claude Fable 5's flagship xhigh configuration on pass@1 — 68.5% versus 69.9%. That gap is basically a rounding error next to Kimi's price tag: $4.65 per rollout against Fable's $13.41.

Give both models a couple more swings and the ranking actually flips. At pass@2, Kimi edges ahead 82.0% to 80.2%. At pass@4, it's 89.4% versus 88.5%, good enough to sit near the top of the entire 44-config DeepSWE export, trailing only two GPT configurations. The reason comes down to coverage versus reliability: Kimi eventually solves 89.4% of all 113 tasks at least once, more than any model tested, but it's less consistent, going four-for-four on only 45 tasks compared to Fable's 58. Fable is the steady hand; Kimi is the wide net that keeps almost everything within reach given enough attempts.

The economics are where this gets hard to ignore. Across the full 452-rollout run, Kimi cost $2,103 total against Fable's $6,010, and per solved task Kimi delivers 14.7 solves per $100 versus Fable's 5.3 — nearly triple the throughput per dollar. For any team running agentic coding workflows at volume, where retries are cheap and expected, that math starts to matter more than a 1.4-point accuracy gap.

What's genuinely odd is how similar these two models are under the hood. Their per-task correlation is 0.72, the highest cross-vendor similarity Together AI has measured on this benchmark, and there isn't a single task where one model aces it four-for-four while the other completely whiffs. They fail in the same ways too — about 65% of failures for both are near misses, and both largely avoid breaking existing tests. Pairing them wouldn't buy much diversity; their combined coverage is 105 of 113 tasks, barely above what Kimi manages alone.

Broken down by language, Fable still leads Python, JavaScript, TypeScript, and Rust, while Kimi takes Go outright and, notably, closes the gap on Rust further than any other model tested, including GPT-5.6 Sol. None of that changes the headline, though: an open-weight model just showed up within striking distance of Anthropic's best coding configuration, at roughly a third of the cost, with better multi-attempt performance to boot.

My take

This is exactly the pattern I expect to keep repeating through 2026 — open-weight models catching frontier closed models within a rounding error, then winning on pass@k and cost the moment you allow more than one try. If your workflow tolerates retries, which most agentic coding pipelines do, paying triple for Fable's marginal reliability edge is starting to look like brand loyalty, not engineering.

Read more about this at: Together AI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.