NVIDIA Measures 49% More AI Output Under a Fixed Power Budget at Nscale
DSX MaxLPS let the test run an additional inference job, while the 99th-percentile wait for the first token rose 17%.
Listen to this story
The audio brief
Story brief
3 key pointsNVIDIA’s Nscale test suggests power pooling can raise AI-factory throughput by fitting another inference job into an unchanged site reservation, rather than making each job dramatically faster. On GB300 systems, DSX MaxLPS lifted aggregate output 49.2%, while high-throughput instance performance barely moved. The tradeoff was a 17% increase in 99th-percentile time to first token, and actual power use rose as the...
- 01
Both configurations had a 264.4 kW provisioned budget; utilization rose from 62.9% to 75.2% with MaxLPS.
- 02
The managed fleet grew from 140 to 192 GPUs, primarily by adding a third high-throughput job.
- 03
High-throughput instance output changed only from 59,153 to 59,220 tokens per second; low-latency output remained 2,265.
Sharing unused power across GPUs let NVIDIA and Nscale run more AI inference work within the same reserved budget. In their newly published evaluation, DSX MaxLPS managed 192 GPUs instead of 140 and delivered 49.2% more aggregate throughput. The gain came with a higher 99th-percentile wait for the first token.
More jobs, nearly unchanged output per job
The test ran Kimi K2.5 inference on NVIDIA GB300 NVL72 systems at Nscale’s Verne campus in Iceland. Its static setup assigned 140 GPUs to two jobs tuned for high throughput and one tuned for low latency. The MaxLPS setup added a third high-throughput job, bringing the managed fleet to 192 GPUs while leaving the low-latency job at the same size.
NVIDIA measured total output rising from 1,084,503 to 1,618,443 tokens per second. The increase chiefly reflects the additional job, rather than a substantial speedup for each existing one. Output per high-throughput instance moved from 59,153 to 59,220 tokens per second; the low-latency instance remained at 2,265.
The budget stayed fixed; consumption did not
Static planning reserves enough power for equipment to reach its specified peak. When one node draws less, that unused portion of its reservation cannot automatically help another. NVIDIA says DSX MaxLPS instead monitors consumption and adjusts GPU power limits across a group of machines. Operators set the group’s budget, priorities and reserves; the software reallocates available power within those rules. It does not increase the site’s supply.
Both configurations had a 264.4 kW provisioned budget, but neither used all of it. Budget utilization increased from 62.9% to 75.2% as the larger fleet did more work. The result is greater output from the same reserved capacity, not the same output with less electricity consumed.
A slower first token at the tail
The throughput gain did not leave every response-time measure untouched. Median and 75th-percentile latency stayed within 5% of the static baseline. But the 99th-percentile time to first token rose 17% from a 15.7-second baseline. That percentile is the wait at or below which 99% of requests began receiving output; it does not measure the absolute slowest request. A service judged by its slow-response threshold could weigh that change differently from one judged mainly by total output.
NVIDIA had previously described capacity gains for MaxLPS in a different Vera Rubin setting. The Nscale evaluation gives a measured result on GB300 NVL72 hardware instead; its 49.2% throughput gain should not be treated as a Vera Rubin result or as a promise for every workload.
What an operator would have to prove
The opportunity depends on workloads leaving power available at different times. Jobs that peak together offer less room to share. NVIDIA also cautions that missing or delayed power measurements can undermine the control decisions that keep a group within its limit. At Nscale, site-level telemetry was used to check rack-level measurements.
NVIDIA advises operators to start with a representative baseline, add capacity in stages and check performance, reserves and power compliance before setting production limits. That makes the Nscale result a useful test case, not a ready-made operating setting for other facilities.
Editorial analysis
Our Read
The distinction worth keeping in view is between a fixed power allocation and fixed electricity use. Nscale’s test shows a way to turn reserved but unused capacity into more output; it does not show that the added work comes without an energy or response-time cost. Our view is that the next persuasive evidence would be a staged production evaluation across changing workloads, including periods when jobs demand power at the same time. That would show whether the measured gain survives the conditions that give operators less headroom to share.
Citation desk / original work
Cite this
Citation desk / original work
Cite this
The distinction worth keeping in view is between a fixed power allocation and fixed electricity use.
/posts/nvidia-measures-49-more-ai-output-under-a-fixed-power-budget-at-nscale#finding-1
Sources
- developer.nvidia.comHow NVIDIA Groq 3 LPX Deterministic Execution Drives Power-Efficient High-Interactivity Inference on NVIDIA Vera Rubin | NVIDIA Technical Blog
- developer.nvidia.comHow NVIDIA DSX MaxLPS Maximizes AI Factory Throughput and Efficiency | NVIDIA Technical Blog
Reader comments
Newest comments first. Replies stay oldest first.