NvidiaZhihan Jiang4 min readintermediate
NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut
Summary
NVIDIA’s Vera Rubin NVL72 AI inference system shows up to 3.7× higher throughput than the prior GB300 NVL72 on MLPerf v6.1 benchmarks (Qwen3‑VL, DeepSeek‑R1), achieves 99% scaling efficiency across 288 GPUs, and benefits from software optimizations (NVFP4 precision, kernel fusion, disaggregated serving). The post is a product announcement with concrete benchmark numbers but limited technical dept…
- Vera Rubin NVL72 delivers up to 3.7× (Qwen3‑VL) and 2.5× (DeepSeek‑R1) higher throughput versus GB300 NVL72 in MLPerf v6.1 offline, server, and interactive scenarios.
- Scaling tests on GB300 NVL72 (4 × 72‑GPU racks) show 99% scaling efficiency, with near‑linear throughput growth.
- Software improvements (lower KV‑cache precision, kernel fusion, disaggregated serving with vLLM/Dynamo) add up to 1.6× performance over MLPerf v6.0.
- Hardware co‑design features—enhanced Tensor Cores, Transformer Engine, NVFP4 precision, 6th‑gen NVLink with 10× packet rate—are credited for gains.
For engineers evaluating AI inference infrastructure, the reported throughput and scaling gains illustrate the impact of tight hardware‑software co‑design on large‑scale LLM serving. The numbers provide a reference point for capacity planning and cost‑per‑token calculations when comparing NVIDIA’s…
5/10





