TLDRocket
Sign in

Moonshot opens Kimi K3 weights — but few can run it

The New Stack Amanda Caswell Covered by 33 sources

Moonshot just open-sourced Kimi K3, a massive coding-focused AI model, after demand crushed its API. Only problem: the weights alone are 1.4 terabytes, so basically nobody but big labs can actually run it.

Moonshot AI dropped the open weights for Kimi K3 on Hugging Face this Monday, and the timing tells you something. The company had just paused new API signups because demand outstripped its GPU capacity within 48 hours. So instead of just scaling up servers, Moonshot handed the model to anyone willing to try their own infrastructure.

The catch is that "anyone" is a very small club. K3 is a 2.8-trillion-parameter Mixture-of-Experts model shipped in the MXFP4 format, and the weights alone take up roughly 1.4 terabytes of storage. Running it for real means a distributed setup of eight or more servers, each packing eight Nvidia H100 or B200 GPUs. That's not a homelab project. That's a line item in a corporate budget with its own procurement cycle.

Moonshot built K3 with an OpenAI-compatible API, which lowers the switching cost for teams already wired into GPT or Claude-style integrations — swap the endpoint, keep the code. Paired with a one-million-token context window, the pitch is clearly aimed at enterprise coding agents and long-horizon knowledge work, not chatbot users. Early testing backs some of that up: The New Stack's own head-to-head with Anthropic's Fable 5 found K3 competitive on coding tasks at about a third of the API cost, though it ran roughly four times slower. Developers are already using it for gnarly systems work, including porting the Godot game engine to WebGPU.

But the benchmark numbers deserve some skepticism, and Moonshot itself seems to know that — its documentation openly flags quirks like K3 defaulting to maximum reasoning effort and occasionally overreaching on ambiguous prompts. Morningstar analyst Malik Ahmed Khan put it plainly: progress, yes, but not parity with American frontier models on real tasks. Then there's the geopolitical noise. Anthropic and White House officials have accused Moonshot of distilling outputs from U.S. models during training, which Moonshot denies, pointing instead to architectural tweaks like Kimi Delta Attention. Whether or not that holds up, it adds a compliance question mark that enterprise buyers now have to weigh alongside the GPU bill.

The real story here isn't whether K3 beats GPT-5.6 Sol on a leaderboard. It's that owning your model is becoming a genuine strategic option, not just an ideological preference — especially after Anthropic's Fable 5 got yanked offline by a Commerce Department directive. Open weights mean nothing if you can't afford to run them, and Moonshot just proved that scale cuts both ways: too much demand breaks your API, too much model breaks everyone's hardware budget.

My take

I keep hearing that open weights mean freedom, but 1.4TB and 64 H100s is not freedom for 99% of developers — it's freedom for hyperscalers and a handful of well-funded labs, full stop. The distillation accusations are a sideshow; the real signal is that access-versus-ownership debate getting sharper every time a U.S. regulator can flip a switch on a rented model. I'd rather see more mid-sized open models nobody needs a data center to run than another trillion-parameter flex nobody can actually deploy.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.