TLDRocket
Sign in

Moonshot AI releases Kimi K3 open model with 2.8 trillion parameters and 1-million-token context window

Open source release Confirmed 92% confidence first seen

Moonshot AI released Kimi K3, a 2.8-trillion-parameter sparse mixture-of-experts model featuring native vision capabilities, a 1-million-token context window, and novel attention mechanisms. The model achieves significant efficiency improvements including 6.3x faster decoding in long contexts and 21% reduced output token usage compared to its predecessor. Full weights are scheduled for release by July 27, 2026, with API pricing starting at $0.30 per million input tokens.

Decision brief

What changed
Moonshot AI released Kimi K3, a 2.8-trillion-parameter sparse mixture-of-experts open model with native vision, a 1-million-token context window, and new attention mechanisms (Kimi Delta Attention, Attention Residuals), claiming 6.3x faster long-context decoding and a 21% reduction in output tokens versus its K2.6 predecessor; full weights are slated for July 27, 2026, with API pricing starting at $0.30 per million input tokens.
Why it matters
For teams running long-context or agentic workloads (e.g., code, chip design, scientific research), the claimed token-efficiency and speed gains could materially cut inference cost and latency if they hold up outside Moonshot's own benchmarks. The release also triggered a rushed, benchmark-free counter-announcement from Alibaba (Qwen3.8-Max), signaling intensifying open-weight competition among Chinese labs that could accelerate price/feature pressure on incumbent model providers.
Affected roles
CTO CFO CISO
Evidence
Specs and architecture details are corroborated across MarkTechPost and The Neuron, with token-efficiency figures separately attributed to Artificial Analysis; the competitive reaction (Alibaba's Qwen3.8-Max preview) is reported consistently by MarkTechPost, The Neuron, and The New Stack, with the latter noting Alibaba provided no benchmarks or model card, unlike Moonshot's fuller disclosure.
What remains uncertain
Full model weights have not yet shipped (promised by July 27, 2026), so real-world reproducibility of the 6.3x decoding speedup and 21% token reduction outside Moonshot's own evaluations is unverified; how K3 compares to proprietary frontier models on independent, third-party benchmarks is also not established in this coverage.
Monitor next
Watch whether Moonshot delivers the full open-weight release on July 27, 2026, and whether independent benchmarking (e.g., from Artificial Analysis or others) confirms the claimed efficiency and performance gains.

Analytical support, not advice — assumptions and open questions stated above.

Source coverage

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.