Moonshot AI Releases Kimi K3: A 2.8 Trillion Parameter Open MoE Model With Kimi Delta Attention and 1M Context
MarkTechPost Asif Razzaq ● Covered by 5 sources
Moonshot AI released Kimi K3, a 2.8-trillion-parameter sparse mixture-of-experts model with native vision and 1-million-token context window, featuring novel attention mechanisms called Kimi Delta Attention and Attention Residuals. The model achieves 6.3x faster decoding in million-token contexts and 25% higher training efficiency, while activating only 16 of 896 experts through Stable LatentMoE sparsity. K3 outperforms other open models on Moonshot's evaluations but remains behind proprietary models like Claude and GPT variants, expanding the scope of openly available large language models.
Why it matters
Moonshot AI released Kimi K3 on July 16, 2026. It is a 2.8-trillion-parameter open MoE model built on Kimi Delta Attention and Attention Residuals, activating 16 of 896 experts. The post Moonshot AI Releases Kimi K3: A 2.8 Trillion Parameter Open MoE Model With Kimi Delta Attention and 1M Context appeared first on MarkTechPost.
Also covered by
- The Neuron — Alibaba previews 2.4 trillion parameter Qwen3.8-Max model for open release
- MarkTechPost — Alibaba Previews Qwen3.8-Max, a 2.4 Trillion-Parameter Multimodal Model, Days After Moonshot’s Kimi K3 Open-Weight Launch
- The Neuron — Artificial Analysis Reports Kimi K3 Token Efficiency
- The Neuron — Moonshot's Kimi K3 Open Model Released