NVIDIA Vera Rubin NVL72 Achieves Up to 3.7x Performance Over GB300 in MLPerf Inference Benchmark Debut
Key point
NVIDIA Vera Rubin NVL72 took the lead in MLPerf Inference v6.1, recording up to 3.7x the throughput of the GB300.
Details
NVIDIA's next-generation platform Vera Rubin NVL72 made its debut in the MLPerf Inference v6.1 benchmark, achieving up to 3.7x higher throughput compared to the previous generation GB300 NVL72. This result demonstrates system performance and infrastructure scaling efficiency, which are key to the economics of AI inference.
Hardware and Software Full-Stack Optimization
The performance improvement of the Vera Rubin NVL72 stems from the tight integration of hardware and software (full-stack codesign). Enhanced Tensor Cores and the Transformer Engine accelerate the prefill and decode stages of inference, while the application of NVFP4 precision reduces the memory footprint, increasing throughput without loss of output quality. Additionally, it leverages Disaggregated serving (separation of prefill/decode) and large-scale expert parallelism for MoE layers. In particular, the NVL72 scale-up domain based on 6th-generation NVLink and NVLink Switch increases packet transmission rates by 10x and reduces latency by 3x compared to standard Ethernet.
Model-Specific Performance and Agent Workloads
In the detailed benchmark results, the Qwen3-VL model showed up to 3.7x the throughput of the GB300 across all offline, server, and interactive scenarios when using vLLM and the NVIDIA Dynamo framework. The DeepSeek-R1 model recorded up to 2.5x performance improvement when using the TensorRT-LLM library. Regarding agent workloads, it demonstrated 30x better performance than the GB300 in the SemiAnalysis AgentX benchmark preview test.
Scaling Efficiency and Continuous Software Improvements
The GB300 NVL72 achieved 99% scaling efficiency in a 4-rack (288-GPU) configuration, showing that throughput increases almost linearly with the number of GPUs. This is the result of high-bandwidth interconnects within racks and efficient request orchestration between nodes. Furthermore, NVIDIA recorded up to 1.6x performance improvement in v6.1 submissions compared to v6.0, and continues to improve performance through additional software optimizations such as lower KV cache precision and kernel fusion even after the v6.1 deadline.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.