MLPerf v6.1 GPT-OSS-120B per-GPU throughput leadership
Company: CoreWeave
The claim, verbatim
CoreWeave achieved the highest per-GPU throughput of any MLPerf v6.1 Datacenter Closed submission for GPT-OSS-120B reasoning model, sustaining 16,635 tokens per second per GPU in offline scenario and 16,118 in server scenario
Source (primary)
MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave (-, news_article)
View cached copy (2026-09-19)Live source ↗
Quote: “CoreWeave sustained 16,635 tokens per second per GPU in offline scenario, and 16,118 in server scenario. These were the highest per-GPU throughput of any v6.1 Datacenter Closed submission on GPT-OSS-120B, on any silicon, in both scenarios.”
How we checked this
Checked on September 24, 2026. The cited source supports every part of this claim.
The blog gives both per-GPU figures and says they were the highest of any v6.1 Datacenter Closed GPT-OSS-120B submission, on any silicon. The post footnotes per-GPU throughput as a derived metric not verified by MLCommons.
Confirmed in the source:
- CoreWeave sustained 16,635 tokens per second per GPU on GPT-OSS-120B in the offline scenario
- CoreWeave sustained 16,118 tokens per second per GPU on GPT-OSS-120B in the server scenario
- These were the highest per-GPU throughput of any MLPerf Inference v6.1 Datacenter Closed submission on GPT-OSS-120B, in both scenarios
- GPT-OSS-120B is described as a reasoning model
What we did: Read our cached copy of the publisher (https://www.coreweave.com/blog/coreweave-leads-cloud-providers-in-mlperf-r-inference-v6-1-performance-with-nvidia-blackwell-ultra) in full (14,189 characters, retrieved September 19, 2026) and checked each assertion in the claim against it.
Additional evidence
confirms MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave
Quote: “CoreWeave sustained 16,635 tokens per second per GPU in offline scenario, and 16,118 in server scenario. These were the highest per-GPU throughput of any v6.1 Datacenter Closed submission on GPT-OSS-120B, on any silicon, in both scenarios.”
