TLDRocket
Sign in

Kimi K3, and what we can still learn from the pelican benchmark

Simon Willison's Weblog Simon Willison Covered by 7 sources

Moonshot AI dropped Kimi K3, a 2.8 trillion parameter model that beats a lot of the big names on benchmarks. It's also the priciest Chinese model yet, and Simon Willison's pelican test shows why cost matters.

Based on reporting by Simon Willison's Weblog, Simon Willison — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Moonshot AI rolled out Kimi K3 today, calling it their most capable model ever at 2.8 trillion parameters. It's live on their website and API now, with open weights promised sometime by July 27, 2026. Moonshot is branding this the first "open 3T-class model," a title it grabs from DeepSeek's 1.6T v4 Pro, and the self-reported benchmarks have it beating Claude Opus 4.8 max and GPT-5.5 high in most categories, though it still trails Claude Fable 5 and GPT-5.6 Sol.

The numbers from Artificial Analysis back up the hype in places. K3 hits an Elo of 1547 on their long-horizon knowledge work test, a jump of 732 points over Kimi K2.6, second only to Claude Fable 5. It's also using 21% fewer output tokens than its predecessor and now leads Arena.ai's Frontend Code arena, ahead of even Fable 5. That's real progress for a lab that was, not long ago, mostly known for undercutting Western pricing.

Which makes the pricing shift here so notable. K3 costs $3 per million input tokens and $15 per million output tokens, putting it right alongside Anthropic's Sonnet line and making it the most expensive release from any Chinese lab to date. Compare that to Kimi K2.6's $0.95/$4, and you can see Moonshot is betting that quality now justifies charging like the Western labs it's chasing.

Willison, as usual, ran his pelican-riding-a-bicycle test on it, and the result is telling in ways the benchmarks aren't. The prompt burned through 16,658 output tokens — 13,241 of them pure reasoning — for a single SVG, costing 25 cents total. K3 apparently only offers one reasoning effort level, "max," and it shows: there's no dial to turn down the thinking, so even a trivial task gets the full expensive treatment. Vision performed well when Willison fed the image back in for alt text, generating an accurate, detailed description for under a cent.

Willison is also candid that the pelican test has lost most of its predictive power over the past 21 months — GLM-5.2 now draws better pelicans than either GPT-5.6 or Claude Fable 5, despite clearly not being in their class overall. What the test still does well is force him to actually run the model, catch quirks like Kimi's mysterious 85-token hidden system prompt, and give a rough sense of cost and reasoning overhead before anyone commits real money to a workload.

My take — AI-written commentary, not fact-checked reporting

A model that can't dial down its reasoning effort is going to burn cash on every trivial request, and 25 cents for an SVG of a bird on a bike is not a small tell. Moonshot wants to charge Western prices now that they've closed the capability gap, fine, but pricing parity should come with control parity too — give me a cheap mode. The real story buried in here isn't the benchmark wins, it's that Chinese labs racing toward parameter counts and price tags that mirror OpenAI and Anthropic is exactly the opposite of what made this ecosystem interesting a year ago.

Read more about this at: Simon Willison's Weblog

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.