[AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing
Latent Space ● Covered by 33 sources
Moonshot AI just dropped Kimi K3, a 2.8-trillion-parameter open-weight model, with full weights due July 27. Early rankings put it near Opus 4.8 and even ahead of Claude on frontend coding — at Sonnet-level prices.
Based on reporting by Latent Space — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Z.ai's GLM had been hogging the open-model spotlight for a while, and Moonshot AI just answered with something enormous. Kimi K3 lands at 2.8 trillion total parameters, a 1M-token context window, and native multimodal input, with Moonshot promising open weights by July 27, 2026. If that ships as described, several people tracking the space are already calling it the largest open-weight model ever released.
The headline numbers are eye-catching on their own, but the architecture is the more interesting story for anyone who actually builds with these things. K3 pairs something called Kimi Delta Attention, or KDA, with a technique Moonshot dubs Attention Residuals. Moonshot says KDA delivers up to 6.3x faster decoding at million-token context lengths, while Attention Residuals adds roughly 25% more training efficiency for under 2% extra cost. Community digging into the release also surfaced a sparse mixture-of-experts setup — 16 experts activated out of 896 — putting the activation ratio under 2%, plus a new activation function called SiTU. One engineer noted KDA's design work reportedly started back in January 2025, meaning it took about a year and a half to scale into a frontier-sized model.
On independent evaluation, Artificial Analysis scored K3 at 57 on its Intelligence Index, putting it in the same neighborhood as Opus 4.8 and GPT-5.5, though still behind Claude Fable 5 and GPT-5.6 Sol. Arena's numbers told a punchier story: K3 jumped from #18 to #1 in Frontend Code Arena, posting a 76% pairwise win rate against 63% for Fable 5 and 58% for GPT-5.6 Sol. In Text Arena it landed at #9, up from #38. Not every metric flattered the model — Artificial Analysis flagged that hallucination rate actually worsened to 51% from 39% even as raw accuracy improved, and ProgramBench's author publicly objected to Moonshot using an averaged-implementation metric rather than counting fully working programs, which can make results look rosier than they are.
Pricing is where K3 gets genuinely interesting for people who have to pay the bill. Reported rates come in at $3 per million input tokens and $15 per million output tokens, with a 90% discount on cached input down to $0.30. Several people compared that directly to Sonnet 5, noting Sonnet was briefly cheaper until end of August before prices converge. A blended estimate at an 80/20 input-output split put K3 around $5.40 per million tokens, versus $9 for Opus 4.8 and $10 for GPT-5.5. Artificial Analysis' own per-task cost estimate came out to $0.94, below GPT-5.6 Sol's $1.04 and well under Opus 4.8's $1.80.
But open weights and cheap don't automatically mean easy to run. Moonshot's own blog reportedly recommends supernode configurations with 64 or more accelerators for efficient inference, and early live testing on OpenRouter clocked speeds around 26 to 28 tokens per second — slower than Opus, with one observer guessing speculative decoding simply wasn't turned on yet. vLLM confirmed Moonshot contributed a KDA-specific prefix caching implementation directly into the project, necessary because KDA breaks the assumptions standard prefix caching relies on. That's a genuinely useful contribution to the ecosystem, but it also underlines that serving a 2.8T model well is its own engineering project, not something you casually spin up on a laptop.
My take — AI-written commentary, not fact-checked reporting
Frontier performance at Sonnet-adjacent pricing, released openly, is the kind of thing that should worry anyone betting on closed-model margins holding steady — but the
Read more about this at: Latent Space