Claude Fable 5 vs. Kimi K3: Same results, one-third the cost, 4x slower
The New Stack Nick Lucchesi ● Covered by 7 sources
Moonshot's Kimi K3 matched Anthropic's Claude Fable 5 on three real coding tasks for a third of the price. But it took about four times as long to get there.
Based on reporting by The New Stack, Nick Lucchesi — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Moonshot AI dropped Kimi K3 in mid-July with a big claim: a 2.8-trillion-parameter open-weight model, the largest ever built in China, that goes toe-to-toe with Anthropic's Opus 4.8. The pitch leans hard on price. Kimi K3's API runs $3 per million input tokens and $15 per million output, versus $10 and $50 for Anthropic's Fable 5. That's more than triple the cost on both ends for Fable. A discount that steep only means something if the work holds up, so one tester decided to find out by running both models through three identical coding jobs and tracking every token spent.
The setup used fd, sharkdp's Rust-based file finder, chosen because it's real production code with a documented bug history. Kimi ran inside Moonshot's own Kimi Code terminal agent, version 0.27.0, while Fable ran in Claude Code version 2.1.212. Same prompts, same repo, fresh folders for each of the six runs. Three tasks: a bug fix, a multi-file refactor, and a new feature.
The bug fix produced identical results. Both models found the exact same root cause in fd's ignore-file logic and deleted the exact same line of code, byte for byte matching diffs, all 70 tests passing on each side. Fable finished in 1 minute 4 seconds using about 347,000 tokens for 85 cents. Kimi took 3 minutes 7 seconds, burned fewer tokens at 238,000, and cost six cents. Cheaper, yes. But nowhere close in speed.
The refactor told a similar story with a twist. Both models split the same bloated function into a new module and kept all 264 tests green, landing on diffs of comparable size. Kimi, though, went further by snapshotting the old binary and diffing it against the new one across roughly 40 CLI scenarios to confirm nothing broke. That thoroughness cost time: 14 minutes 50 seconds and 928,000 tokens, the slowest run the tester has clocked in months. Fable wrapped the same job in 3 minutes 11 seconds on 639,000 tokens. Kimi came in at 70 cents against Fable's $2.32.
The feature build, a new --count flag for fd, followed the pattern. Both delivered working code. Kimi touched six files and updated the man page; Fable touched seven and also updated zsh completions, stumbling briefly on a test guess before self-correcting twice. Kimi needed 10 minutes 21 seconds and 2.1 million tokens for $1.38, nearly five times longer than Fable's 2 minutes 34 seconds, roughly 1.46 million tokens, and $2.81.
Add it all up and Kimi finished the three jobs for $2.13 against Fable's $5.98, close to the third-of-the-price promise Moonshot made. But Kimi used more tokens overall, 3.3 million versus 2.4 million, so the savings come purely from cheaper rates, not efficiency. And the real gap shows up in the clock: Fable cleared all three tasks in under 7 minutes, Kimi needed just over 28. For a model chasing established players, that kind of lag is the harder problem to fix.
My take — AI-written commentary, not fact-checked reporting
A price cut that comes bundled with a fourfold time penalty isn't really a discount, it's a different product with a different set of trade-offs. Anyone billing by the hour, or just tired of watching a spinner, will feel that gap long before they notice the invoice. Kimi K3 clearly gets the answers right, which matters, but in coding tools speed is half the product, and Moonshot has a lot of ground to make up there before the price tag becomes the headline instead of the caveat.
Read more about this at: The New Stack