GPT-OSS-120B per-GPU offline throughput
Company: CoreWeave
The claim, verbatim
CoreWeave sustained 16,635 tokens per second per GPU in offline scenario for GPT-OSS-120B, achieving the highest per-GPU throughput of any MLPerf v6.1 Datacenter Closed submission
Source (primary)
MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave (-, news_article)
View cached copy (2026-09-19)Live source ↗
Quote: “CoreWeave sustained 16,635 tokens per second per GPU in offline scenario”
How we checked this
Checked on September 24, 2026. The cited source supports every part of this claim.
The post reports 16,635 tokens per second per GPU offline on GPT-OSS-120B and calls it the highest per-GPU throughput of any v6.1 Datacenter Closed submission on that model. It also notes that per-GPU throughput is a derived metric not verified by MLCommons.
Confirmed in the source:
- CoreWeave sustained 16,635 tokens per second per GPU in the offline scenario on GPT-OSS-120B
- This was the highest per-GPU throughput of any MLPerf v6.1 Datacenter Closed submission on GPT-OSS-120B
What we did: Read our cached copy of the publisher (https://www.coreweave.com/blog/coreweave-leads-cloud-providers-in-mlperf-r-inference-v6-1-performance-with-nvidia-blackwell-ultra) in full (14,189 characters, retrieved September 19, 2026) and checked each assertion in the claim against it.
Additional evidence
confirms MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave
Quote: “CoreWeave sustained 16,635 tokens per second per GPU in offline scenario”
