proomt

Search

Search posts, papers, and topics

All posts

NvidiaZhihan Jiang4 min readintermediate

NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

Summary

NVIDIA’s Vera Rubin NVL72 AI inference system shows up to 3.7× higher throughput than the prior GB300 NVL72 on MLPerf v6.1 benchmarks (Qwen3‑VL, DeepSeek‑R1), achieves 99% scaling efficiency across 288 GPUs, and benefits from software optimizations (NVFP4 precision, kernel fusion, disaggregated serving). The post is a product announcement with concrete benchmark numbers but limited technical dept…

  • Vera Rubin NVL72 delivers up to 3.7× (Qwen3‑VL) and 2.5× (DeepSeek‑R1) higher throughput versus GB300 NVL72 in MLPerf v6.1 offline, server, and interactive scenarios.
  • Scaling tests on GB300 NVL72 (4 × 72‑GPU racks) show 99% scaling efficiency, with near‑linear throughput growth.
  • Software improvements (lower KV‑cache precision, kernel fusion, disaggregated serving with vLLM/Dynamo) add up to 1.6× performance over MLPerf v6.0.
  • Hardware co‑design features—enhanced Tensor Cores, Transformer Engine, NVFP4 precision, 6th‑gen NVLink with 10× packet rate—are credited for gains.

For engineers evaluating AI inference infrastructure, the reported throughput and scaling gains illustrate the impact of tight hardware‑software co‑design on large‑scale LLM serving. The numbers provide a reference point for capacity planning and cost‑per‑token calculations when comparing NVIDIA’s…

5/10

Related reading

  1. From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production

    NVIDIA’s DSX platform lets AI data‑centers shift workloads in response to grid signals, squeezing ~24% more token throughput (4 M→5 M tps) and ~23% better performance‑per‑watt on a fixed megawatt budget. The first production demo used Emerald AI’s Conductor to drop a 4 MW load to 3 MW in under a minute without interrupting high‑priority jobs. DSX MaxLPS reallocates headroom across HGX B200 server…

    Nvidianvidia.com5 min
  2. 5 Companies Using NVIDIA AI for Clean Energy

    Nvidia’s blog spotlights five companies that are using Nvidia AI platforms to accelerate clean‑energy projects—from grid interconnection and nuclear plant operations to off‑grid AI data‑center power, advanced reactors, and fusion tokamaks. The article is a marketing summary and provides few technical details.

    Nvidianvidia.com4 min
  3. Why Deploying Physical AI at Scale Demands Safety at Every Layer

    NVIDIA’s Halos platform is a full‑stack safety system for physical AI (autonomous vehicles and industrial robots). It bundles safety‑engineered hardware (DRIVE AGX Thor, IGX Thor), an ASIL‑D certified OS (Halos OS), middleware for isolation and monitoring, AI models for explainability (Alpamayo), and simulation/validation tools (Isaac Lab, Omniverse). The blog argues that scaling physical AI requ…

    Nvidianvidia.com5 min
  4. ‘Now We Can Know Everything and Do Anything,’ Jensen Huang Says at Dreamforce

    Nvidia’s Jensen Huang announced Salesforce’s Koa, a CRM‑reasoning LLM built by fine‑tuning Nvidia Nemotron 3 Super on a synthetic, three‑decade‑spanning dataset. Koa uses supervised fine‑tuning plus RL (NeMo RL, Gym, AutoModel), covers 14+ industries, and claims 3× fewer errors on Salesforce’s CRM‑Bench versus leading models. It’s already in internal Slack agents and slated for limited customer p…

    Nvidianvidia.com3 min