Nvidia (NVDA) Debuts Vera Rubin NVL72 in MLPerf Inference v6.1

Sarah Jane Sep 17, 2026 5 min read

Nvidia (NVDA) published the first MLPerf results for its Vera Rubin NVL72 platform on Sept. 16, reporting throughput gains of up to 3.7 times its current generation in the MLPerf Inference v6.1 round released the same day by MLCommons.

Nvidia said it submitted Vera Rubin results on two benchmarks: DeepSeek-R1, run with TensorRT-LLM, and Qwen3-VL, run with vLLM using the company’s Dynamo framework. On Qwen3-VL the company reported up to 3.7 times the throughput of its GB300 NVL72 system, and on DeepSeek-R1 up to 2.5 times.

The platform uses what Nvidia describes as enhanced Tensor Cores, an upgraded Transformer Engine with NVFP4 precision, and sixth-generation NVLink with NVLink Switch for interconnect within and across racks.

Nvidia also reported results for the existing GB300 NVL72, saying a 288-GPU configuration spanning four racks achieved 99% scaling efficiency, and that software improvements between v6.0 and v6.1 delivered up to 1.6 times higher performance on the same hardware. A figure of 30 times better performance on agentic inference was described by the company as coming from preview testing rather than a completed benchmark submission.

MLCommons said the round drew a record 30 submitting organisations. Among them were several vendors whose hardware is widely deployed in enterprise data centres, including AMD, Cisco, Dell, Hewlett Packard Enterprise, Intel, Supermicro, Google, Microsoft Azure and Oracle.

The v6.1 round added an end-to-end retrieval-augmented generation benchmark covering embedding models, retrievers, re-rankers and language models, plus an edge agentic inference benchmark measuring multi-turn conversational workloads under fixed memory and power limits. MLCommons said visual language model results improved 2.99 times over v6.0, and DeepSeek-R1 results improved 5.7 times against v5.1 a year earlier.

Buyers specifying accelerated infrastructure can browse our Nvidia GPU and accelerator range, and our guide on GPU and accelerator server requirements covers the power, cooling and PCIe limits these systems impose.

Sources: Nvidia corporate blog, Sept. 16, 2026; MLCommons MLPerf Inference v6.1 results announcement, Sept. 16, 2026.

Sarah Jane

Sarah Jane

Senior IT Hardware Specialist · TechSellerUSA
Sarah helps businesses and IT teams source the right enterprise hardware at wholesale prices. View profile →