TLDRocket
Sign in

To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!

Together AI

Together’s kernels team gained access to the NVIDIA Vera Rubin NVL72 platform and updated ThunderKittens to run NVFP4 and FP8 GEMMs on it by adding Rubin-specific ISA and memory/pipeline features. The NVFP4 GEMM kernels improve from about 42.1% of the roofline to over 22 PFLOPS. These changes increase attainable performance by doubling the K step, expanding tensor memory and shared memory use, and adjusting tiling and commit/collector behaviors so tensor cores stay fed.

Why it matters

We ported ThunderKittens to NVIDIA's Vera Rubin NVL72 and rebuilt our NVFP4 GEMM around the new hardware, taking it from 42% of roofline to over 22 PFLOPS — competitive with cuBLAS and CuTe DSL. Here is what changed in the ISA and how we used it.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.