Moonshot AI Open-Sources MoonEP: A Perfectly Balanced Expert Parallelism Library for MoE Training
MarkTechPost Michal Sutter ● Covered by 29 sources
Moonshot AI open-sourced MoonEP, a library for efficient distributed Mixture-of-Experts training that guarantees every GPU rank receives exactly S × K tokens regardless of router imbalance. The library achieved a claimed 2.5× improvement in scaling efficiency for their Kimi K3 model, a 2.8-trillion-parameter MoE system, by using redundantly planned experts and static buffer shapes to eliminate communication overhead and memory fragmentation. MoonEP's approach removes per-layer host synchronization and makes communication latency nearly immune to router skew, allowing frameworks to scale MoE training more reliably across many GPUs.
Why it matters
Moonshot AI has open-sourced MoonEP, an Expert Parallelism (EP) communication library for distributed Mixture-of-Experts (MoE) workloads. The team announced the release as a library built to make expert-parallel communication more efficient at scale. It ships under an MIT license. MoonEP arrived as part of Kimi K3 Open Day. Alongside the K3 model weights and technical […] The post Moonshot AI Open-Sources MoonEP: A Perfectly Balanced Expert Parallelism Library for MoE Training appeared first on MarkTechPost.
Also covered by
- AWS Machine Learning — Deploying Kimi K3 on AWS
- Rest of World — With Moonshot’s free Kimi K3, China changes the sovereign AI playbook
- Together AI — Together AI announces strategic partnership with Moonshot AI to natively serve Kimi models
- MarkTechPost — Building Non-Interactive Agentic Coding Workflows with Moonshot AI’s Kimi CLI, JSONL Streaming, Testing, and Session Memory
- The Neuron — Moonshot Releases Kimi K3 Model Weights on Hugging Face
- Simon Willison — moonshotai/Kimi-K3
- MarkTechPost — Kimi AI and kvcache-ai Open Sources ‘AgentENV’: A Distributed System that Powers Agentic Reinforcement Learning (RL) Training for Kimi K3
- The New Stack — Moonshot opens Kimi K3 weights — but few can run it
- The Verge — Why China is giving away its best AI models
- TLDR — Silicon Valley Splits Over Closing the Borders to Chinese AI
- TechCrunch AI — Making sense of the panic over Chinese AI
- Together AI — Kimi K3 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing
- TechCrunch AI — ‘AI communism’, rogue models, and the why Kimi K3 spooked Wall Street
- Together AI — Kimi K3 vs Claude Fable 5 on DeepSWE: Cost and Coding
- Exponential View — 🔮 Will Kimi K3 change the economics of AI?
- ChinaTalk — Kimi and Xi
- Interconnects — Open models recap: more on Kimi K3, Qwen 3.8, Xi's WAIC speech, distillation, the open-closed gap, and what's next
- The New Stack — Moonshot launched Kimi K3. Then demand shut down subscriptions in 48 hours.
- Latent Space — [AINews] not much happened today
- Interconnects — Kimi K3: The open-weights escalation
- Zvi (Don't Worry About the Vase) — On Kimi K3: Its Capabilities And Related Discontents
- Exponential View — 📈 Data to start your week
- Import AI — Import AI 465: Open vs closed gaps; Kimi K3; Demis' big policy plan
- The Neuron — Kimi K3 Pushed Open Models Toward the Frontier
- TheSequence — The Sequence Radar #897: Last Week in AI: China, Compression and the Open-Model Race
- Latent Space — [AINews] not much happened today
- The New Stack — Kimi K3 tops Arena’s coding leaderboard — and it’s open-weight
- Latent Space — [AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing