TLDRocket
Sign in

Moonshot's Kimi K3 Open Model Released

Kimi Covered by 7 sources

Moonshot just released Kimi K3, an open model with 2.8 trillion parameters and a 1-million-token context window. It's the largest open model ever built, and it's nipping at the heels of GPT and Claude's best.

Based on reporting by Kimi — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Moonshot AI didn't ease into this one. Kimi K3 lands as a 2.8-trillion-parameter model, the first open release to cross into what the company calls the "3T-class," and it ships today across Kimi.com, Kimi Work, Kimi Code, and the API. Full weights won't drop until July 27, 2026, but the model is already live and thinking at maximum effort by default, with lighter and heavier modes promised in later updates.

The engineering underneath is where things get interesting. Kimi K3 runs on two new architectural pieces, Kimi Delta Attention and Attention Residuals, paired with a sparser Mixture of Experts setup that activates just 16 of 896 experts through something called Stable LatentMoE. Moonshot says this combination, plus refined training recipes, delivers roughly 2.5 times the scaling efficiency of its predecessor Kimi K2. That's the kind of claim that matters more than the headline parameter count, because it's the difference between a model that's big and a model that's usably big.

Where K3 tries to prove itself is in long, messy, agentic work. In one test, it optimized GPU kernels across Nvidia H200 and a rival GPGPU platform for up to 24 hours per task, and reportedly beat GPT 5.6 Sol and Claude Opus 4.8 while running close to Claude Fable 5. In another, it built a Triton-style GPU compiler called MiniTriton from scratch, complete with its own IR layer and PTX code generation, and the thing trained nanoGPT end to end with a stable loss curve. Moonshot even had it design a chip in a single 48-hour autonomous run, an INT4 accelerator with 8,700 tokens per second of simulated decode throughput on a 45nm process. Whether or not that chip ever gets fabricated, it's a decent stunt.

On the knowledge-work side, K3 leans into being multimodal and visually literate. It produced a 42-year ASIC industry report by pulling from 2,800 web searches and 11,000 pages, analyzed 391 gravitational-wave events with 20 concurrent subagents, and edited its own promotional video from 56 raw clips, including beat-matched cuts and audio sync. Moonshot frames all of this as evidence of frontier-level performance, even while admitting the model still trails Claude Fable 5 and GPT 5.6 Sol overall.

Pricing undercuts the big closed labs by a wide margin: 30 cents per million tokens for cache hits, three dollars for cache misses, fifteen for output. Combined with a promised cache-hit rate above 90% on coding workloads via Moonshot's Mooncake infrastructure, K3 is clearly built to be run hard and often, not just admired from a benchmark table.

My take — AI-written commentary, not fact-checked reporting

I'll say the obvious thing nobody wants to: an open 2.8-trillion-parameter model beating Opus 4.8 on real engineering tasks is a bigger deal than another 2-point benchmark bump from a closed lab charging you by the drink. Moonshot keeps proving that open-weight releases from Chinese labs are setting the pace on raw scale while US labs guard weights like state secrets, and Europe just watches from the sidelines building committees instead of models. I don't care that K3 still trails GPT 5.6 Sol on the leaderboard — the fact that anyone can eventually download these weights and run them without a subscription is the actual headline.

Read more about this at: Kimi

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.