Cursor Open-Sources Mixture-of-Kittens (MoK): A Deterministic MoE Training Megakernel for GB300 NVL72 Racks
MarkTechPost 3 weeks ago 44 ● 2 sources
Cursor Research open-sourced Mixture-of-Kittens, a mixture-of-experts training kernel that fuses all MoE communication and computation into a single deterministic megakernel for large GPU clusters.The kernel achieves up to 2.37x higher throughput than existing baselines and requires NVIDIA Blackwell GPUs in GB300 NVL72 racks with Python 3.12+, PyTorch 2.10+, and CUDA 13.0+.Organizations with access to large-scale GPU infrastructure can now use MoK under Apache-2.0 to accelerate training of mixture-of-experts models like DeepSeek-V3-style architectures.