TLDRocket
Sign in

Fable 5.1 vs. Fable 5: Results on a real-world budget, not the spec sheet

The New Stack Jessica Wachtel Covered by 2 sources

Anthropic’s Claude Fable 5.1 posted a higher Terminal-Bench-Science result than Claude Fable 5, but an independent test of the benchmark’s tasks under a fixed user-like budget found a much smaller gap. In the sampled runs with a $12 limit and 60 turns per test, Fable 5.1 solved 1 task where Fable 5 solved 0, with total cost $40.75 vs $53.59 and total time 262 min vs 388 min. The benchmark’s claimed doubling did not translate into a similarly large advantage for average-style runs, and differences mostly showed up in cost and whether the model hit the imposed budget limit.

Why it matters

When Anthropic launched Claude Fable 5.1 this month, it centered the announcement around one benchmark result: its Terminal-Bench-Science score. In The post Fable 5.1 vs. Fable 5: Results on a real-world budget, not the spec sheet appeared first on The New Stack.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.