TLDRocket
Sign in

Custom Kernels for All from Codex and Claude

Hugging Face

Codex and Claude generated production-ready CUDA kernels for a video diffusion pipeline and a language model using a new agent skill that packages GPU optimization knowledge. The RMSNorm kernel achieved 1.88x speedup on isolated benchmarks for LTX-Video and 1.94x for Qwen3-8B, translating to 6% and measurable end-to-end improvements when combined with torch.compile. Users can now install the skill, prompt an agent to build optimized kernels, benchmark them against PyTorch baselines, and publish pre-compiled binaries to the HuggingFace Kernel Hub for one-line loading by others.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.