NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut
NVIDIA Blog Zhihan Jiang ● Covered by 6 sources
NVIDIA says its new Vera Rubin NVL72 tops GB300 NVL72 in MLPerf inference tests. The bigger point: the wins come from hardware, software, and networking all moving together.
Based on reporting by NVIDIA Blog, Zhihan Jiang — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
NVIDIA used the latest MLPerf Inference v6.1 results to make a familiar argument with fresh numbers: AI inference economics depend on raw speed, scaling, and constant software work. In its first preview submission for Vera Rubin NVL72, the system posted leading results on two demanding tests, DeepSeek-R1 and Qwen3-VL. On Qwen3-VL, Vera Rubin NVL72 delivered up to 3.7x higher throughput than GB300 NVL72 across offline, server, and interactive modes. On DeepSeek-R1, the gain was up to 2.5x, using NVIDIA TensorRT-LLM.
The company is pitching those results as more than bragging rights. Higher throughput means more tokens, more users served, and lower cost per token. That is the business story NVIDIA wants people to hear, and it leans heavily on full-stack codesign: enhanced Tensor Cores, Transformer Engine, NVFP4 precision, and disaggregated serving that splits prefill and decode. For mixture-of-experts models like DeepSeek-R1 and Qwen3-VL, that matters because the work is spread across layers that can get messy fast if the plumbing is weak.
NVIDIA also pointed to scale. A DeepSeek-R1 submission on GB300 NVL72 stretched from one rack, or 72 GPUs, to four racks, or 288 GPUs, and the company says it hit 99% scaling efficiency in the offline scenario. Throughput rose nearly in line with the added hardware. On the WAN 2.2 text-to-video benchmark, GB300 NVL72 reached 0.65 720p videos per second at 5.7 seconds per video, which NVIDIA says was 9x higher throughput and 7.5x lower latency than a single node.
Software remains part of the pitch too. NVIDIA says Qwen3-VL performance on GB300 NVL72 improved up to 1.6x over v6.0 thanks to lower KV cache precision, kernel fusion, better kernels, and disaggregated serving with vLLM and NVIDIA Dynamo. Some post-submission results on GPT-OSS-120B and DLRMv3 improved further, though those are not yet verified by MLCommons. And outside the big racks, NVIDIA also submitted Jetson AGX Thor results on the new Edge-Agentic benchmark with Qwen3.6-27B, while 19 partners showed up across the ecosystem, including Nebius, Oracle Cloud Infrastructure, HPE, Supermicro, and others.
My take — AI-written commentary, not fact-checked reporting
This is classic NVIDIA: make the rack, the software, and the benchmark look like one machine, because that’s where the moat lives. The open-source bits are useful, but the real message is that closed-stack integration still prints the best scores. Everybody else gets to call it “competition”; NVIDIA gets to call it Tuesday.
Read more about this at: NVIDIA Blog