Fable 5.1 vs. Fable 5: Results on a real-world budget, not the spec sheet
The New Stack Jessica Wachtel ● Covered by 2 sources
Anthropic’s Claude Fable 5.1 posted a higher Terminal-Bench-Science result than Claude Fable 5, but an independent test of the benchmark’s tasks under a fixed user-like budget found a much smaller gap. In the sampled runs with a $12 limit and 60 turns per test, Fable 5.1 solved 1 task where Fable 5 solved 0, with total cost $40.75 vs $53.59 and total time 262 min vs 388 min. The benchmark’s claimed doubling did not translate into a similarly large advantage for average-style runs, and differences mostly showed up in cost and whether the model hit the imposed budget limit.
Why it matters
When Anthropic launched Claude Fable 5.1 this month, it centered the announcement around one benchmark result: its Terminal-Bench-Science score. In The post Fable 5.1 vs. Fable 5: Results on a real-world budget, not the spec sheet appeared first on The New Stack.