TLDRocket
Sign in

Custom Kernels for All from Codex and Claude

Hugging Face Blog

Codex and Claude generated production-ready CUDA kernels for a video diffusion pipeline and a language model using a new agent skill that packages GPU optimization knowledge. The RMSNorm kernel achieved 1.88x speedup on isolated benchmarks for LTX-Video and 1.94x for Qwen3-8B, translating to 6% and measurable end-to-end improvements when combined with torch.compile. Users can now install the skill, prompt an agent to build optimized kernels, benchmark them against PyTorch baselines, and publish pre-compiled binaries to the HuggingFace Kernel Hub for one-line loading by others.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.