Toolspublished

NVIDIA Says MaxLPS Can Add Up to 40% More Rubin GPUs Within a Fixed Power Budget

The suite treats unused rack headroom, workload tuning and cooling power as capacity to recover. Its biggest capacity claim is still a projection, while its control software remains in preview.

By 2 min read
NVIDIA Says MaxLPS Can Add Up to 40% More Rubin GPUs Within a Fixed Power Budget

Listen to this story

The audio brief

About 1:45
0:001:45
Read transcript
NVIDIA says its DSX MaxLPS suite could fit up to 40% more Vera Rubin NVL72 GPU capacity inside the same facility power budget. The important qualifier is that 40% is a planning projection, not a production deployment result—and the control software is still in Developer Preview. The approach is to recover power that normally sits idle. In NVIDIA’s illustrative 100-megawatt AI factory, only 60 megawatts reaches the AI load after facility overhead, rack losses, and operational inefficiency. MaxLPS lets operators set power budgets for groups of racks, nodes, or GPUs, then move unused headroom to active systems without exceeding the site limit. In representative inference tests, NVIDIA says provisioned power fell from 125 to 90 kilowatts for a GB200 NVL72 rack, and from 136 to 101 kilowatts for a Vera Rubin NVL72 rack. It reports roughly 1.5 times the performance per watt for GB200, and 1.3 to 1.4 times for Rubin. Those settings reportedly supported 39% more GB200 racks and 35% more Rubin racks within the same power envelope, while maintaining workload throughput. The software also tunes GPUs for inference, training, memory-bound, or compute-bound work. And Rubin racks are designed for 45-degree-Celsius liquid-cooling inlets, which could increase free cooling depending on climate and site design. The constraint to watch is whether these representative tests—and the projected 40% gain—translate into operating deployments at scale.

Story brief

3 key points

NVIDIA is pitching DSX MaxLPS as a way to increase accelerator capacity without expanding a data center’s power envelope. In representative inference tests, the software and workload-specific settings reportedly enabled 39% more GB200 racks and 35% more Vera Rubin racks, while cutting provisioned rack power from 125 to 90 kW and 136 to 101 kW, respectively. The headline projection is up to 40% more Rubin NVL72...

  1. 01

    NVIDIA’s 100 MW illustration allocates 60 MW to AI load after facility overhead, rack losses, and operational inefficiency.

  2. 02

    Dynamic Power Software lets operators set group budgets and reassign unused capacity across racks, nodes, and GPUs.

  3. 03

    NVIDIA reports 1.5x performance per watt for GB200 and 1.3x–1.4x for Rubin in representative inference configurations.

NVIDIA says DSX MaxLPS can raise AI-factory output within a fixed facility power budget by shifting unused rack power to active GPUs, tailoring GPU settings to workloads and reducing cooling overhead.

The immediate prize is more installed hardware. NVIDIA projects that MaxLPS, combined with data-center power planning, can support up to 40% more Vera Rubin NVL72 GPU capacity under the same power budget. That figure is a planning projection, not a disclosed deployment result.

The power lost before it reaches compute

Grid input is not the same as power available to AI chips. In NVIDIA’s illustrative 100 MW factory, 20 MW goes to facility overhead, 10 MW to rack losses and 10 MW to operational inefficiency, leaving 60 MW for AI load.

Traditional planning reserves each rack’s maximum expected draw, even when workloads consume less. MaxLPS is designed to make that headroom available elsewhere while keeping the site power envelope unchanged.

Representative inference rack power

GB200 NVL72
125 kW90 kW
kW

NVIDIA says MaxLPS reduced provisioned rack power in a representative GB200 NVL72 inference test.

Vera Rubin NVL72
136 kW101 kW
kW

NVIDIA reports the corresponding reduction for Vera Rubin NVL72 in representative inference testing.

A control loop with limits set by operators

Dynamic Power Software monitors use from the utility level through racks, nodes and GPUs. Operators define resource groups, budgets and allocation policies; the software can then reassign unused capacity within those managed groups.

The software collects telemetry, identifies headroom, reallocates power within policy and checks compliance with group budgets. It is currently in Developer Preview; NVIDIA’s disclosed validation centers on representative inference workloads.

Software settings meet facility design

MaxLPS includes workload profiles for inference, training, memory-bound and compute-bound jobs. NVIDIA’s Application Performance and Power Manager applies the selected GPU configuration instead of using one operating point for every workload.

Vera Rubin NVL72 racks are designed for 45°C liquid-cooling inlet operation. NVIDIA says warmer coolant can increase free cooling and reduce reliance on mechanical chillers, depending on climate and site design.

In those inference tests, NVIDIA reports about 1.5x performance per watt on GB200 NVL72 and 1.3x to 1.4x on Vera Rubin NVL72. It says the configurations enabled 39% more GB200 racks and 35% more Rubin racks within the same power envelope while preserving workload throughput.

NVIDIA’s longer-term model is to size power, cooling, space and network infrastructure for an eventual rack count, then populate additional positions as workload mix shifts and average rack power declines. That strategy requires the infrastructure to be sized for the lifecycle target up front.

Sources

  1. developer.nvidia.comMaximizing AI Factory Performance per Watt with NVIDIA DSX MaxLPS | NVIDIA Technical Blog