CoreWeave Publishes MLPerf Results, Claims Cloud-Provider Leads

Its GB300 systems led selected cloud-provider comparisons, CoreWeave says, while its broader GPT-OSS per-GPU ranking uses a company-derived metric that MLCommons did not verify.

By 2 min read
CoreWeave Publishes MLPerf Results, Claims Cloud-Provider Leads
CoreWeave Publishes MLPerf Results, Claims Cloud-Provider Leads

Listen to this story

The audio brief

About 1:49
0:001:49
Read transcript
CoreWeave says its NVIDIA GB300 NVL72 systems delivered the strongest results among cloud providers in selected MLPerf tests—and that the numbers came from production infrastructure, not a benchmark-only setup. In the latest MLPerf Inference results, CoreWeave entered four NVIDIA platforms across four model families. On the GB300 NVL72, a rack containing 72 GPUs, it reported 1,196 queries per second for Qwen3-VL-235B-A22B. For Llama 2 70B, it reported 944,902 tokens per second in the server test and 1,136,100 in the offline test. CoreWeave says those were the top cloud-provider results using that GB300 system. The broader claim needs more care. For GPT-OSS-120B, CoreWeave calculated 16,635 offline tokens per second per GPU, and 16,118 in the server scenario, calling both the highest per-GPU figures in the full submission set. But those per-GPU numbers are the company’s own normalization: MLCommons verified the underlying submission, not the derived ranking. CoreWeave also reported a 19.8 percent improvement in per-GPU server throughput for DeepSeek-R1-671B versus its earlier GB200 NVL72 result, while noting that the two tests used different GPU counts. The strategic point is production readiness: CoreWeave says customer-available images, clusters, and operational software produced the scores. The key constraint now is comparability—whether other providers publish equivalent per-GPU calculations and whether those calculations receive the same scrutiny.

Story brief

3 key points

CoreWeave is using MLPerf Inference v6.1 to make a production-readiness case for its NVIDIA GB300 NVL72 cloud infrastructure, not just a benchmark win. Its reported scores include 1,196 queries per second on Qwen3-VL-235B-A22B and 944,902 server tokens per second on Llama 2 70B. The strongest GPT-OSS-120B rankings are derived per-GPU figures that MLCommons did not verify, so customers should distinguish certified...

  1. 01

    CoreWeave entered four NVIDIA platforms and four model families in MLPerf’s Datacenter Closed, Available category.

  2. 02

    The company reported 16,635 offline tokens per second per GPU for GPT-OSS-120B on one GB300 NVL72 rack.

  3. 03

    CoreWeave says production clusters, customer-available images, and operational software—not benchmark-only tuning—powered the submissions.

CoreWeave has published MLPerf Inference v6.1 results that it says put its NVIDIA GB300 NVL72 systems ahead of other cloud-provider submissions on selected multimodal and language-model workloads. The company is also arguing that those results came from production clusters, not a system tuned only for a benchmark.

CoreWeave entered four NVIDIA platforms and four model families in MLPerf’s Datacenter Closed, Available category. Its strongest provider-specific claims concern the GB300 NVL72, a 72-GPU rack system: the company reported 1,196 queries a second on Qwen3-VL-235B-A22B’s server test, which it called the highest among cloud providers.

For Llama 2 70B, CoreWeave reported 944,902 tokens a second in the server scenario and 1,136,100 in the offline scenario. It described both as the highest throughput among cloud-provider submissions using GB300 NVL72.

The GPT-OSS-120B claim is broader than the provider comparisons. CoreWeave reported 16,635 tokens a second per GPU offline and 16,118 in the server scenario from one GB300 NVL72 rack, calling both the highest per-GPU throughput of any v6.1 Datacenter Closed submission.

That ranking carries a material qualification. CoreWeave calculates per-GPU throughput by dividing a system’s total throughput by its number of accelerators. It says MLCommons did not verify that derived metric, so the per-GPU ranking has a different verification status from the underlying benchmark submission.

CoreWeave also reported a 19.8% increase in derived per-GPU server throughput for DeepSeek-R1-671B against its earlier MLPerf v6.0 GB200 NVL72 result. The submissions used different GPU counts—72 in v6.1 and 64 in v6.0—so CoreWeave normalized the comparison per GPU rather than presenting it as a total-system score.

CoreWeave says the submissions used production clusters, images and operational software available to customers, rather than a benchmark-specific tuned cluster. That makes the company’s case about more than a leaderboard position: it is presenting the measurements as evidence that its operating stack can deliver those results outside a dedicated test environment.

Sources

  1. coreweave.comMLPerf® Inference v6.1 Results: CoreWeave Leads Providers

Loading discussion...