GPT-OSS-120B aggregate throughput on GB300 NVL72 offline
Company: CoreWeave
The claim, verbatim
CoreWeave achieved 1.19 million tokens per second in offline scenario on a single NVIDIA GB300 NVL72 rack for GPT-OSS-120B
Source (primary)
MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave (-, news_article)
View cached copy (2026-09-19)Live source ↗
Quote: “over 1.19 million tokens per second in offline scenarios on a single NVIDIA GB300 NVL72 rack”
How we checked this
Checked on September 24, 2026. The cited source supports every part of this claim.
The post states CoreWeave achieved over 1.19 million tokens per second in the offline scenario on a single GB300 NVL72 rack for GPT-OSS-120B.
Confirmed in the source:
- CoreWeave achieved over 1.19 million tokens per second in the offline scenario on GPT-OSS-120B
- The result was on a single NVIDIA GB300 NVL72 rack
What we did: Read our cached copy of the publisher (https://www.coreweave.com/blog/coreweave-leads-cloud-providers-in-mlperf-r-inference-v6-1-performance-with-nvidia-blackwell-ultra) in full (14,189 characters, retrieved September 19, 2026) and checked each assertion in the claim against it.
Additional evidence
confirms MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave
Quote: “over 1.19 million tokens per second in offline scenarios on a single NVIDIA GB300 NVL72 rack”
