NVIDIA Puts Power Management at Center of Its AI Infrastructure Push
At its AI Infra Summit, NVIDIA paired a measured Lambda result with larger company-reported Vera Rubin benchmarks and projections, arguing that electricity allocation—not just chip speed—will govern AI capacity.
Listen to this story
The audio brief
Story brief
3 key pointsNVIDIA is positioning power-aware infrastructure as a differentiator for AI factories, pairing workload-level controls with grid-responsive operations and resilient networking. Lambda’s fixed-budget test put 19 Blackwell nodes where 16 full-power nodes would normally fit, raising token throughput 24%. NVIDIA’s larger Vera Rubin claims—up to 40% more GPU capacity per megawatt and 30× throughput per megawatt versus...
- 01
Lambda’s test increased throughput from roughly 4 million to 5 million tokens per second and improved performance per watt 23%.
- 02
DSX MaxLPS reallocates unused power across GPUs and racks; NVIDIA claims up to 1.4× more tokens per megawatt.
- 03
DSX Flex and Emerald AI’s Conductor can throttle or pause lower-priority jobs when grid or price signals require load reduction.
NVIDIA used its AI Infra Summit to argue that the decisive measure for large AI systems is no longer a chip’s peak speed, but how many useful AI tokens a facility produces from each megawatt. Its case rested on two kinds of evidence: a Lambda test that fit more Blackwell servers into a fixed power budget, and NVIDIA’s larger Vera Rubin performance claims.
Lambda said it ran 19 Blackwell server nodes on a budget normally allocated to 16 full-power nodes. It reported cluster token throughput rose 24%, from about 4 million to 5 million tokens a second, while performance per watt improved 23%. NVIDIA presented that as the first validation of its DSX MaxLPS power-management software on Blackwell servers.
One validation, larger platform claims
MaxLPS continuously monitors consumption across GPUs and racks, then shifts available power toward workloads that need it and recovers capacity left unused by static provisioning. NVIDIA says this is particularly useful where training and inference share a facility but draw power differently. The company says MaxLPS can deliver up to 1.4 times more tokens per megawatt through factory-wide optimization.
The more ambitious figures have different footing. NVIDIA says suitable Vera Rubin NVL72 environments could support up to 40% more GPU capacity within the same megawatt budget. Separately, it reported up to 30 times higher throughput per megawatt than GB300 NVL72 on DeepSeek V4 Pro in SemiAnalysis AgentX results, alongside up to 45 times lower cost per million tokens. Those are NVIDIA-reported results for a specified model and workload, not a general guarantee for every site.
A data center that responds to the grid
NVIDIA’s other argument is that an AI facility can be a flexible electricity customer rather than a fixed load. With Silicon Valley Power, Emerald AI and NVIDIA demonstrated automated load reduction in response to hundreds of grid demand signals while protecting priority workloads. The proposed operating model is explicit: lower-priority AI jobs can be throttled or paused temporarily, then resume when conditions allow.
Emerald AI plans to use NVIDIA DSX Flex with its Conductor software to adjust a facility’s energy use according to real-time grid signals and hybrid energy sources. The software can receive load-shedding requests, demand-response events and price signals within a workload hierarchy. Capacity for critical work is preserved by accepting interruptions to work deemed less urgent.
Keeping the system running is part of the case
NVIDIA also tied energy efficiency to keeping large systems running. It described NVLink 6 as using fault-detection, containment and recovery features at physical and network layers, aimed at preventing local failures from becoming wider stalls. NVIDIA says the sixth-generation NVLink provides 3.6 TB/s of bidirectional bandwidth per GPU and 260 TB/s across a 72-GPU Vera Rubin NVL72 rack domain.
NVIDIA is selling an AI-factory design in which software directs electricity, grid signals can reshape lower-priority computing, and networking protects output from a power-constrained installation. Lambda’s result is a concrete proof point; the broader Vera Rubin claims will depend on how they translate across more environments.
Sources
- developer.nvidia.comNVIDIA NVLink: The Scale-Up Network for AI Factories | NVIDIA Technical Blog
- blogs.nvidia.comAI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories
Loading discussion...
Reader comments
Newest comments first. Replies stay oldest first.