NVIDIA Says MaxLPS Can Add Up to 40% More Rubin GPUs Within a Fixed Power Budget
The suite treats unused rack headroom, workload tuning and cooling power as capacity to recover. Its biggest capacity claim is still a projection, while its control software remains in preview.
Listen to this story
The audio brief
Story brief
3 key pointsNVIDIA is pitching DSX MaxLPS as a way to increase accelerator capacity without expanding a data center’s power envelope. In representative inference tests, the software and workload-specific settings reportedly enabled 39% more GB200 racks and 35% more Vera Rubin racks, while cutting provisioned rack power from 125 to 90 kW and 136 to 101 kW, respectively. The headline projection is up to 40% more Rubin NVL72...
- 01
NVIDIA’s 100 MW illustration allocates 60 MW to AI load after facility overhead, rack losses, and operational inefficiency.
- 02
Dynamic Power Software lets operators set group budgets and reassign unused capacity across racks, nodes, and GPUs.
- 03
NVIDIA reports 1.5x performance per watt for GB200 and 1.3x–1.4x for Rubin in representative inference configurations.
NVIDIA says DSX MaxLPS can raise AI-factory output within a fixed facility power budget by shifting unused rack power to active GPUs, tailoring GPU settings to workloads and reducing cooling overhead.
The immediate prize is more installed hardware. NVIDIA projects that MaxLPS, combined with data-center power planning, can support up to 40% more Vera Rubin NVL72 GPU capacity under the same power budget. That figure is a planning projection, not a disclosed deployment result.
The power lost before it reaches compute
Grid input is not the same as power available to AI chips. In NVIDIA’s illustrative 100 MW factory, 20 MW goes to facility overhead, 10 MW to rack losses and 10 MW to operational inefficiency, leaving 60 MW for AI load.
Traditional planning reserves each rack’s maximum expected draw, even when workloads consume less. MaxLPS is designed to make that headroom available elsewhere while keeping the site power envelope unchanged.
Representative inference rack power
NVIDIA says MaxLPS reduced provisioned rack power in a representative GB200 NVL72 inference test.
NVIDIA reports the corresponding reduction for Vera Rubin NVL72 in representative inference testing.
A control loop with limits set by operators
Dynamic Power Software monitors use from the utility level through racks, nodes and GPUs. Operators define resource groups, budgets and allocation policies; the software can then reassign unused capacity within those managed groups.
The software collects telemetry, identifies headroom, reallocates power within policy and checks compliance with group budgets. It is currently in Developer Preview; NVIDIA’s disclosed validation centers on representative inference workloads.
Software settings meet facility design
MaxLPS includes workload profiles for inference, training, memory-bound and compute-bound jobs. NVIDIA’s Application Performance and Power Manager applies the selected GPU configuration instead of using one operating point for every workload.
Vera Rubin NVL72 racks are designed for 45°C liquid-cooling inlet operation. NVIDIA says warmer coolant can increase free cooling and reduce reliance on mechanical chillers, depending on climate and site design.
In those inference tests, NVIDIA reports about 1.5x performance per watt on GB200 NVL72 and 1.3x to 1.4x on Vera Rubin NVL72. It says the configurations enabled 39% more GB200 racks and 35% more Rubin racks within the same power envelope while preserving workload throughput.
NVIDIA’s longer-term model is to size power, cooling, space and network infrastructure for an eventual rack count, then populate additional positions as workload mix shifts and average rack power declines. That strategy requires the infrastructure to be sized for the lifecycle target up front.
Sources
- developer.nvidia.comMaximizing AI Factory Performance per Watt with NVIDIA DSX MaxLPS | NVIDIA Technical Blog