GPT-OSS-120B on GB200 NVL72 server
Company: CoreWeave
The claim, verbatim
CoreWeave's NVIDIA GB200 NVL72 achieved 901,058 tokens per second in server scenario for GPT-OSS-120B, the highest rack-scale total in its class
Source (primary)
MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave (-, news_article)
View cached copy (2026-09-19)Live source ↗
Quote: “Our NVIDIA GB200 NVL72 posted the highest rack-scale GPT-OSS-120B total in its class, reaching 901,058 tokens per second in server”
How we checked this
Checked on September 24, 2026. The cited source supports every part of this claim.
The post says CoreWeave's GB200 NVL72 posted the highest rack-scale GPT-OSS-120B total in its class, reaching 901,058 tokens per second in the server scenario.
Confirmed in the source:
- CoreWeave's NVIDIA GB200 NVL72 reached 901,058 tokens per second in the server scenario on GPT-OSS-120B
- This was the highest rack-scale GPT-OSS-120B total in its class
What we did: Read our cached copy of the publisher (https://www.coreweave.com/blog/coreweave-leads-cloud-providers-in-mlperf-r-inference-v6-1-performance-with-nvidia-blackwell-ultra) in full (14,189 characters, retrieved September 19, 2026) and checked each assertion in the claim against it.
Additional evidence
confirms MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave
Quote: “Our NVIDIA GB200 NVL72 posted the highest rack-scale GPT-OSS-120B total in its class, reaching 901,058 tokens per second in server”
