TLDRocket
Sign in

Claude Fable 5.1 made me a really nice animated pelican

Simon Willison’s Weblog Simon Willison Covered by 3 sources

Claude Fable 5.1 is out, and it can draw a very good animated pelican on a bike. The weird part: the best version came from max reasoning, but low and medium barely reasoned at all.

Based on reporting by Simon Willison’s Weblog, Simon Willison — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Anthropic has rolled out Claude Fable 5.1, alongside Mythos 5.1, and it’s making a bold claim: this is a new bar for coding, knowledge work, and long-running problem solving. The headline number in the launch post is a 52.6% score on the fresh Terminal-Bench-Science 0.1 benchmark, up sharply from 24.7% for Fable 5, 29.0% for Opus 5, and 22.4% for GPT-5.6 Sol. Other benchmark gains are there too, but they don’t hit nearly as hard.

Simon Willison’s real test was simpler and more amusing: “Generate an SVG of a pelican riding a bicycle.” He’s been losing faith in the pelican benchmark as a broad signal, and now mostly treats it as a way to compare models within the same family, especially across reasoning levels. Fable 5.1 has five of those: low, medium, high, xhigh, and max. There’s no way to turn reasoning off entirely.

At low and medium, the model looked oddly flat. The transcripts showed no summarized reasoning text, and the outputs were almost the same size: 1,998 tokens at low, 1,977 at medium. Both ran for about 23 seconds and cost just over 10 cents. High added a little more thought and a slightly more detailed build, but it still didn’t change the result much.

Then things got expensive. Xhigh jumped to 36,767 output tokens, took 7 minutes 51 seconds, and cost $1.83. Max went even further: 65,927 output tokens, 13 minutes 54 seconds, and $3.30. That was also where Willison got the best pelican he’s seen from Anthropic’s models, complete with a tasteful background, feet on the pedals, a wing on the handlebars, a blue hat, and even a basket with a fish.

The funny twist is that the max run spent a lot of time arguing with itself about tiny visual choices — helmet versus crest, feather edges, fork angle, whether to add a bell. Willison even took that max pelican, fed it back into high with “animate this,” and got a video version for $1.37. The wheels apparently spun the wrong way in MP4, which is exactly the sort of thing that makes modern AI feel both impressive and mildly annoying.

My take — AI-written commentary, not fact-checked reporting

This is the sort of benchmark story that tells on the model more than the marketing deck does. If the best pelican only shows up when the system is allowed to think for 14 minutes and burn $3.30, that’s not magic — that’s a very expensive bicycle with a fish basket. Also, “no reasoning” reasoning at low and medium is a nice reminder that these products still enjoy a bit of drama in the plumbing.

Read more about this at: Simon Willison’s Weblog

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.