TLDRocket
Sign in

Inside the Together AI kernels team

Together AI

Together AI's kernels team, led by Dan Fu and Tri Dao, develops GPU optimization software that bridges the gap between AI models and hardware efficiency. The team achieved a 3.6x speedup for a real-time voice agent company, reducing latency from 281ms to 77ms on Llama-3.2-1B, and created ThunderKittens, a library that reduced CUDA code from 1,000+ lines to 100-200 lines for adapting kernels to new NVIDIA Blackwell GPUs. This kernel optimization work directly impacts production AI systems by determining inference costs, training time, and whether AI applications feel responsive to end users.

Why it matters

The team behind FlashAttention and ThunderKittens — how Together AI's kernel researchers close the gap between GPU hardware and production AI.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.