Kimi K3, and what we can still learn from the pelican benchmark
Simon Willison Simon Willison ● Covered by 5 sources
Moonshot AI released Kimi K3, a 2.8 trillion parameter model available via API with open weights promised by July 27, 2026, positioning it as the first open 3-trillion parameter model. The model costs $3 per million input tokens and $15 per million output tokens, making it the most expensive Chinese AI lab model to date and comparable to Anthropic's Claude Sonnet pricing. The author demonstrates K3's capabilities through a pelican-riding-a-bicycle benchmark test, which generates a 16,658-token response costing 25 cents, while reflecting on how this once-useful comparison metric has diminished in correlation with actual model quality as capabilities have advanced.
Why it matters
Chinese AI lab Moonshot AI announced Kimi K3 this morning, describing it as their "most capable model to date, with 2.8 trillion parameters". It's currently available via their website and API, but an open weight release is promised "by July 27, 2026". Moonshot are calling this the first "open 3T-class model" (I guess they're rounding 2.8 trillion up to 3 trillion), taking the crown from DeepSeek's 1.6T v4 Pro. Their self-reported benchmarks have K3 mostly beating Claude Opus 4.8 max and GPT-5.5 high, while losing out to Claude Fable 5 and GPT-5.6 Sol. A few highlights from the Artificial Analysis report on the model: "On our private long-horizon knowledge work evaluation, Kimi K3 reaches an overall Elo of 1547, +732 points from Kimi K2.6 and behind only Claude Fable 5." "Cost per task ($0.94) is similar to GPT-5.6 Sol ($1.04), ~1/2 the price of Opus 4.8 ($1.80) and higher than open weights peers" "Kimi K3’s token usage on the Artificial Analysis Intelligence Index decreased significan
Also covered by
- The New Stack — Claude Fable 5 vs. Kimi K3: Same results, one-third the cost, 4x slower
- TLDR Dev — The Kimi K3 Moment
- Exponential View — 🔮 Kimi K3 surprise & AI economics; the solar paradox; AI's right to learn, cancer vaccine & junior jobs++
- MarkTechPost — Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving Cost