TLDRocket
Sign in

The Pelican comparison grid for Astra is pretty interesting

Simon Willison’s Weblog Simon Willison Covered by 8 sources

Opinion — commentary, not a factual news event.

Astra’s pelican-on-a-bike SVGs beat GPT-5.6 Sol, even at low reasoning. The odd part: Astra used fewer tokens, so the better image could cost less.

Based on reporting by Simon Willison’s Weblog, Simon Willison — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Simon Willison got access to GPT-6 Astra and did what any sensible person would do: asked it for SVGs of pelicans riding bicycles. He ran the prompt across low, medium, high, xhigh and max reasoning levels, then compared the results in a grid against GPT-5.6 Sol, Terra and Luna. The grid turned out to be more than a joke. It made the differences between the models easy to see.

The big takeaway is simple: Astra’s pelicans are better. Willison says the strongest GPT-5.6 Sol version he tried — xhigh, and even that beat max in his view — still looked like abstract shapes. Every Astra pelican from low through xhigh looked better than that, and the max version was described as really good. Astra below max still had a recurring problem, though: it did not reliably place the pelican legs on both sides of the frame.

Cost adds another wrinkle. Astra may be roughly twice as expensive as Sol at $10 per million input tokens and $50 per million output tokens, versus $5 and $30 for Sol. But Astra also used fewer tokens at each level, which narrows the gap. In practice, that means the sticker price does not tell the whole story.

The punchline is hard to ignore. Astra low produced a better pelican than any of the GPT-5.6 Sol models at any level, and did it for 9.55 cents. Willison points out that spending about 10 cents on other models gave a much worse result. He also noticed something curious in the token counts: Astra and Luna both used 16 input tokens, while Sol and Terra used 26. That made him wonder whether Astra and Luna are more closely related than OpenAI has said.

My take — AI-written commentary, not fact-checked reporting

This is the kind of benchmark that matters because it is concrete and slightly ridiculous, which is usually where the truth hides. Fancy model names and reasoning tiers are cheap talk; a pelican on a bike exposes whether the system can actually hold a shape together. OpenAI can keep the branding fireworks, but the grid says the real competition is still on cost, tokens, and whether the legs end up on the right side of the frame.

Read more about this at: Simon Willison’s Weblog

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.