To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!
Together AI's kernels team said it received access to the NVIDIA Vera Rubin NVL72 platform and ported its ThunderKittens kernels to it, adding functionality to write NVFP4 and FP8 GEMMs on Vera Rubin. The team said its Blackwell NVFP4 and FP8 kernels achieved only around 42.1% and 44.4% of the roofline when run naively on Vera Rubin, and that rebuilding the NVFP4 GEMM around the new hardware took it to over 22 PFLOPS, which it described as competitive with cuBLAS and CuTe DSL.
The team attributed the gap to the Blackwell kernel not feeding tensor cores fast enough, and described several Vera Rubin features it used. The K step per MMA can rise from 32 bytes on Blackwell to 64 bytes, expressed through a new template parameter. Tensor memory grows from 512 to 576 columns, reachable only through an .exclusive qualifier added in PTX 9.4. Shared memory can be increased to 328 KiB via a host-side attribute. Vera Rubin also extends the collector buffer to the B tile, using FILL, USE, LASTUSE and DISCARD labels that the team said are permission qualifiers for reuse, not guarantees. PTX 9.4 adds tcgen05.commit.sync_restrict::shared::read::mma::a for earlier A release.
The team reported sweeping 16k square GEMMs across shared memory pipeline depths: for NVFP4, three stages with 202 KiB reached 17,054 TFLOPS, four stages with 258 KiB reached 20,595, and five stages with 314 KiB reached 22,239; for FP8 E4M3, 209 KiB reached 10,895, 257 KiB reached 11,995, and 305 KiB reached 11,288. It measured the B-side collector at roughly 1 to 3 percent improvement, and early A release at 13.5 percent and 22.1 percent speedups for 64k and 128k square NVFP4 GEMMs. L2 eviction hints helped by a few tenths of a percent.
The team said all measurements used NVIDIA CUDA 13.4 on a Qualification Sample GPU, and that it expects baselines to improve with Vera Rubin software releases. Together AI said its kernels and performance teams are hiring.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
We ported ThunderKittens to NVIDIA's Vera Rubin NVL72 and rebuilt our NVFP4 GEMM around the new hardware, taking it from 42% of roofline to over 22 PFLOPS — competitive with cuBLAS and CuTe DSL. Here is what changed in the ISA and how we used it.