Moonshot AI releases Kimi K3 open model with 2.8 trillion parameters and 1-million-token context window
Open source release ● Confirmed 92% confidence first seen
Moonshot AI released Kimi K3, a 2.8-trillion-parameter sparse mixture-of-experts model featuring native vision capabilities, a 1-million-token context window, and novel attention mechanisms. The model achieves significant efficiency improvements including 6.3x faster decoding in long contexts and 21% reduced output token usage compared to its predecessor. Full weights are scheduled for release by July 27, 2026, with API pricing starting at $0.30 per million input tokens.
Decision brief
- What changed
- Moonshot AI released Kimi K3, a 2.8-trillion-parameter sparse mixture-of-experts open model with native vision, a 1-million-token context window, and new attention mechanisms (Kimi Delta Attention, Attention Residuals), claiming 6.3x faster long-context decoding and a 21% reduction in output tokens versus its K2.6 predecessor; full weights are slated for July 27, 2026, with API pricing starting at $0.30 per million input tokens.
- Why it matters
- For teams running long-context or agentic workloads (e.g., code, chip design, scientific research), the claimed token-efficiency and speed gains could materially cut inference cost and latency if they hold up outside Moonshot's own benchmarks. The release also triggered a rushed, benchmark-free counter-announcement from Alibaba (Qwen3.8-Max), signaling intensifying open-weight competition among Chinese labs that could accelerate price/feature pressure on incumbent model providers.
- Evidence
- Specs and architecture details are corroborated across MarkTechPost and The Neuron, with token-efficiency figures separately attributed to Artificial Analysis; the competitive reaction (Alibaba's Qwen3.8-Max preview) is reported consistently by MarkTechPost, The Neuron, and The New Stack, with the latter noting Alibaba provided no benchmarks or model card, unlike Moonshot's fuller disclosure.
- What remains uncertain
- Full model weights have not yet shipped (promised by July 27, 2026), so real-world reproducibility of the 6.3x decoding speedup and 21% token reduction outside Moonshot's own evaluations is unverified; how K3 compares to proprietary frontier models on independent, third-party benchmarks is also not established in this coverage.
- Monitor next
- Watch whether Moonshot delivers the full open-weight release on July 27, 2026, and whether independent benchmarking (e.g., from Artificial Analysis or others) confirms the claimed efficiency and performance gains.
Analytical support, not advice — assumptions and open questions stated above.