NvidiaVishal Ganeriwala5 min readintermediate
From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production
Summary
NVIDIA’s DSX platform lets AI data‑centers shift workloads in response to grid signals, squeezing ~24% more token throughput (4 M→5 M tps) and ~23% better performance‑per‑watt on a fixed megawatt budget. The first production demo used Emerald AI’s Conductor to drop a 4 MW load to 3 MW in under a minute without interrupting high‑priority jobs. DSX MaxLPS reallocates headroom across HGX B200 server…
- Dynamic workload orchestration (Emerald AI Conductor) can respond to utility demand‑response signals in <1 min, preserving high‑priority inference while shedding lower‑priority jobs.
- Lambda’s five‑rack, 19‑node HGX B200 cluster ran 24% more token throughput (4 M→5 M tps) and 23% higher performance‑per‑watt under the same power budget using DSX MaxLPS.
- DSX MaxLPS reallocates GPU and rack‑level power headroom based on workload type, recovering capacity that static provisioning leaves idle.
- NVIDIA’s upcoming 800 VDC power architecture aims to reduce conversion losses and enable denser rack designs for future AI factories.
Power is the primary bottleneck for scaling AI inference and training. By turning AI farms into flexible grid resources and intelligently reallocating power at runtime, NVIDIA shows a path to increase compute density without new transmission infrastructure—critical for meeting the exploding demand…
5/10





