GPT-OSS-120B per-GPU server throughput
Company: CoreWeave
The claim, verbatim
CoreWeave sustained 16,118 tokens per second per GPU in server scenario for GPT-OSS-120B
Source (primary)
MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave (-, news_article)
View cached copy (2026-09-19)Live source ↗
Quote: “16,118 in server scenario”
How we checked this
Checked on September 24, 2026. The cited source supports every part of this claim.
The post reports 16,118 tokens per second per GPU in the server scenario for GPT-OSS-120B on a single GB300 NVL72 rack.
Confirmed in the source:
- CoreWeave sustained 16,118 tokens per second per GPU in the server scenario on GPT-OSS-120B
What we did: Read our cached copy of the publisher (https://www.coreweave.com/blog/coreweave-leads-cloud-providers-in-mlperf-r-inference-v6-1-performance-with-nvidia-blackwell-ultra) in full (14,189 characters, retrieved September 19, 2026) and checked each assertion in the claim against it.
Additional evidence
confirms MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave
Quote: “16,118 in server scenario”
