NVIDIA Measures 49% More AI Output Under a Fixed Power Budget at Nscale

DSX MaxLPS let the test run an additional inference job, while the 99th-percentile wait for the first token rose 17%.

By 3 min read
NVIDIA Measures 49% More AI Output Under a Fixed Power Budget at Nscale
NVIDIA Measures 49% More AI Output Under a Fixed Power Budget at Nscale

Listen to this story

The audio brief

About 1:38
0:001:38
Read transcript
NVIDIA and Nscale measured 49.2 percent more AI inference output under the same 264.4-kilowatt power reservation—not by making existing jobs much faster, but by fitting in another one. In the test, Kimi K2.5 ran on GB300 NVL72 systems at Nscale’s Iceland site. With static power planning, the setup managed 140 GPUs. Using NVIDIA’s DSX MaxLPS software, it managed 192, adding a third high-throughput job. Total output rose from about 1.08 million to 1.62 million tokens per second. But output from each high-throughput job barely changed, and the low-latency job’s output stayed the same. The gain came chiefly from doing more work at once. The power reservation stayed fixed, but actual use went up: utilization rose from 62.9 to 75.2 percent. So this is more output from reserved capacity, not the same output with less electricity. There was a latency tradeoff at the slow end. Median and 75th-percentile latency stayed within five percent of the baseline, while the 99th-percentile wait for the first token increased 17 percent, from 15.7 seconds. That matters for services with strict limits on slow responses. The result depends on jobs having power headroom at different times, and on reliable power measurements. The key question for operators is whether their own workload mix and telemetry can support the extra capacity without breaching latency or power limits.

Story brief

3 key points

NVIDIA’s Nscale test suggests power pooling can raise AI-factory throughput by fitting another inference job into an unchanged site reservation, rather than making each job dramatically faster. On GB300 systems, DSX MaxLPS lifted aggregate output 49.2%, while high-throughput instance performance barely moved. The tradeoff was a 17% increase in 99th-percentile time to first token, and actual power use rose as the...

  1. 01

    Both configurations had a 264.4 kW provisioned budget; utilization rose from 62.9% to 75.2% with MaxLPS.

  2. 02

    The managed fleet grew from 140 to 192 GPUs, primarily by adding a third high-throughput job.

  3. 03

    High-throughput instance output changed only from 59,153 to 59,220 tokens per second; low-latency output remained 2,265.

Sharing unused power across GPUs let NVIDIA and Nscale run more AI inference work within the same reserved budget. In their newly published evaluation, DSX MaxLPS managed 192 GPUs instead of 140 and delivered 49.2% more aggregate throughput. The gain came with a higher 99th-percentile wait for the first token.

More jobs, nearly unchanged output per job

The test ran Kimi K2.5 inference on NVIDIA GB300 NVL72 systems at Nscale’s Verne campus in Iceland. Its static setup assigned 140 GPUs to two jobs tuned for high throughput and one tuned for low latency. The MaxLPS setup added a third high-throughput job, bringing the managed fleet to 192 GPUs while leaving the low-latency job at the same size.

NVIDIA measured total output rising from 1,084,503 to 1,618,443 tokens per second. The increase chiefly reflects the additional job, rather than a substantial speedup for each existing one. Output per high-throughput instance moved from 59,153 to 59,220 tokens per second; the low-latency instance remained at 2,265.

The budget stayed fixed; consumption did not

Static planning reserves enough power for equipment to reach its specified peak. When one node draws less, that unused portion of its reservation cannot automatically help another. NVIDIA says DSX MaxLPS instead monitors consumption and adjusts GPU power limits across a group of machines. Operators set the group’s budget, priorities and reserves; the software reallocates available power within those rules. It does not increase the site’s supply.

Both configurations had a 264.4 kW provisioned budget, but neither used all of it. Budget utilization increased from 62.9% to 75.2% as the larger fleet did more work. The result is greater output from the same reserved capacity, not the same output with less electricity consumed.

A slower first token at the tail

The throughput gain did not leave every response-time measure untouched. Median and 75th-percentile latency stayed within 5% of the static baseline. But the 99th-percentile time to first token rose 17% from a 15.7-second baseline. That percentile is the wait at or below which 99% of requests began receiving output; it does not measure the absolute slowest request. A service judged by its slow-response threshold could weigh that change differently from one judged mainly by total output.

NVIDIA had previously described capacity gains for MaxLPS in a different Vera Rubin setting. The Nscale evaluation gives a measured result on GB300 NVL72 hardware instead; its 49.2% throughput gain should not be treated as a Vera Rubin result or as a promise for every workload.

What an operator would have to prove

The opportunity depends on workloads leaving power available at different times. Jobs that peak together offer less room to share. NVIDIA also cautions that missing or delayed power measurements can undermine the control decisions that keep a group within its limit. At Nscale, site-level telemetry was used to check rack-level measurements.

NVIDIA advises operators to start with a representative baseline, add capacity in stages and check performance, reserves and power compliance before setting production limits. That makes the Nscale result a useful test case, not a ready-made operating setting for other facilities.

Editorial analysis

Our Read

The distinction worth keeping in view is between a fixed power allocation and fixed electricity use. Nscale’s test shows a way to turn reserved but unused capacity into more output; it does not show that the added work comes without an energy or response-time cost. Our view is that the next persuasive evidence would be a staged production evaluation across changing workloads, including periods when jobs demand power at the same time. That would show whether the measured gain survives the conditions that give operators less headroom to share.

Citation desk / original work

Cite this

Permanent attributionView citation
Finding 01

The distinction worth keeping in view is between a fixed power allocation and fixed electricity use.

/posts/nvidia-measures-49-more-ai-output-under-a-fixed-power-budget-at-nscale#finding-1

Sources

  1. developer.nvidia.comHow NVIDIA Groq 3 LPX Deterministic Execution Drives Power-Efficient High-Interactivity Inference on NVIDIA Vera Rubin | NVIDIA Technical Blog
  2. developer.nvidia.comHow NVIDIA DSX MaxLPS Maximizes AI Factory Throughput and Efficiency | NVIDIA Technical Blog

Loading discussion...

YOUR READING SPACE

Notifications