NVIDIA Puts Power Management at Center of Its AI Infrastructure Push

At its AI Infra Summit, NVIDIA paired a measured Lambda result with larger company-reported Vera Rubin benchmarks and projections, arguing that electricity allocation—not just chip speed—will govern AI capacity.

By 3 min read
NVIDIA Puts Power Management at Center of Its AI Infrastructure Push
NVIDIA Puts Power Management at Center of Its AI Infrastructure Push

Listen to this story

The audio brief

About 1:40
0:001:40
Read transcript
Lambda fit 19 Blackwell server nodes into the power budget normally used for 16, and reported that the cluster’s token throughput rose from roughly four million to five million tokens per second. Performance per watt improved 23 percent. That is the concrete proof point behind NVIDIA’s broader argument: for AI infrastructure, the scarce resource may be electricity allocation, not peak chip speed. NVIDIA says its DSX MaxLPS software continuously watches power use across GPUs and racks, shifting available capacity toward workloads that need it and reclaiming power left idle by static provisioning. The company claims that factory-wide optimization can produce up to 1.4 times more tokens per megawatt, particularly when training and inference share a site. The larger claims come from NVIDIA’s own Vera Rubin projections and workload results. It says an NVL72 system could fit up to 40 percent more GPU capacity into the same megawatt budget. Separately, it reported up to 30 times the throughput per megawatt of GB300 NVL72 on DeepSeek V4 Pro, based on SemiAnalysis AgentX results. Those figures are workload-specific, not guarantees for every data center. The operating model also includes the grid. DSX Flex, paired with Emerald AI’s Conductor, can throttle or pause lower-priority jobs when grid demand or prices call for less load. And NVLink 6 is designed to contain failures while connecting 72 GPUs at up to 260 terabytes per second. The key question is whether Lambda’s measured gain—and the bigger Vera Rubin claims—hold across more environments.

Story brief

3 key points

NVIDIA is positioning power-aware infrastructure as a differentiator for AI factories, pairing workload-level controls with grid-responsive operations and resilient networking. Lambda’s fixed-budget test put 19 Blackwell nodes where 16 full-power nodes would normally fit, raising token throughput 24%. NVIDIA’s larger Vera Rubin claims—up to 40% more GPU capacity per megawatt and 30× throughput per megawatt versus...

  1. 01

    Lambda’s test increased throughput from roughly 4 million to 5 million tokens per second and improved performance per watt 23%.

  2. 02

    DSX MaxLPS reallocates unused power across GPUs and racks; NVIDIA claims up to 1.4× more tokens per megawatt.

  3. 03

    DSX Flex and Emerald AI’s Conductor can throttle or pause lower-priority jobs when grid or price signals require load reduction.

NVIDIA used its AI Infra Summit to argue that the decisive measure for large AI systems is no longer a chip’s peak speed, but how many useful AI tokens a facility produces from each megawatt. Its case rested on two kinds of evidence: a Lambda test that fit more Blackwell servers into a fixed power budget, and NVIDIA’s larger Vera Rubin performance claims.

Lambda said it ran 19 Blackwell server nodes on a budget normally allocated to 16 full-power nodes. It reported cluster token throughput rose 24%, from about 4 million to 5 million tokens a second, while performance per watt improved 23%. NVIDIA presented that as the first validation of its DSX MaxLPS power-management software on Blackwell servers.

One validation, larger platform claims

MaxLPS continuously monitors consumption across GPUs and racks, then shifts available power toward workloads that need it and recovers capacity left unused by static provisioning. NVIDIA says this is particularly useful where training and inference share a facility but draw power differently. The company says MaxLPS can deliver up to 1.4 times more tokens per megawatt through factory-wide optimization.

The more ambitious figures have different footing. NVIDIA says suitable Vera Rubin NVL72 environments could support up to 40% more GPU capacity within the same megawatt budget. Separately, it reported up to 30 times higher throughput per megawatt than GB300 NVL72 on DeepSeek V4 Pro in SemiAnalysis AgentX results, alongside up to 45 times lower cost per million tokens. Those are NVIDIA-reported results for a specified model and workload, not a general guarantee for every site.

A data center that responds to the grid

NVIDIA’s other argument is that an AI facility can be a flexible electricity customer rather than a fixed load. With Silicon Valley Power, Emerald AI and NVIDIA demonstrated automated load reduction in response to hundreds of grid demand signals while protecting priority workloads. The proposed operating model is explicit: lower-priority AI jobs can be throttled or paused temporarily, then resume when conditions allow.

Emerald AI plans to use NVIDIA DSX Flex with its Conductor software to adjust a facility’s energy use according to real-time grid signals and hybrid energy sources. The software can receive load-shedding requests, demand-response events and price signals within a workload hierarchy. Capacity for critical work is preserved by accepting interruptions to work deemed less urgent.

Keeping the system running is part of the case

NVIDIA also tied energy efficiency to keeping large systems running. It described NVLink 6 as using fault-detection, containment and recovery features at physical and network layers, aimed at preventing local failures from becoming wider stalls. NVIDIA says the sixth-generation NVLink provides 3.6 TB/s of bidirectional bandwidth per GPU and 260 TB/s across a 72-GPU Vera Rubin NVL72 rack domain.

NVIDIA is selling an AI-factory design in which software directs electricity, grid signals can reshape lower-priority computing, and networking protects output from a power-constrained installation. Lambda’s result is a concrete proof point; the broader Vera Rubin claims will depend on how they translate across more environments.

Sources

  1. developer.nvidia.comNVIDIA NVLink: The Scale-Up Network for AI Factories | NVIDIA Technical Blog
  2. blogs.nvidia.comAI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories

Loading discussion...