To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!
Together AI
Together’s kernels team gained access to the NVIDIA Vera Rubin NVL72 platform and updated ThunderKittens to run NVFP4 and FP8 GEMMs on it by adding Rubin-specific ISA and memory/pipeline features. The NVFP4 GEMM kernels improve from about 42.1% of the roofline to over 22 PFLOPS. These changes increase attainable performance by doubling the K step, expanding tensor memory and shared memory use, and adjusting tiling and commit/collector behaviors so tensor cores stay fed.
Why it matters
We ported ThunderKittens to NVIDIA's Vera Rubin NVL72 and rebuilt our NVFP4 GEMM around the new hardware, taking it from 42% of roofline to over 22 PFLOPS — competitive with cuBLAS and CuTe DSL. Here is what changed in the ISA and how we used it.