[AINews] Megakernels are so dead and so back
Latent Space 5 hours ago 44 ● 2 sources
Megakernels—hand-fused GPU kernels designed to reduce launch overhead—are largely abandoned in production systems due to complexity and the superiority of modular approaches, though research teams continue exploring them; NVIDIA's Rubin GPU includes dependency triggers that further reduce their justification, yet Cursor's open-source Mixture of Kittens megakernel achieved a 41% increase in tokens per second, suggesting the debate remains live. The key technical shift is that NVIDIA's new hardware features make kernel fusion less necessary for hiding latency, while companies like Meta and inference providers have moved toward TensorRT-LLM and modular kernels that optimize individual components and parallelize better. As a result, the megakernel research direction appears to be closing while serving infrastructure consolidates around composable, easier-to-optimize building blocks rather than monolithic fused kernels.