jain.com

GPT-OSS-120B aggregate throughput on GB300 NVL72 server

Company: CoreWeave

Subject kind
product
Statement date
2026-09-16
Promised amount
1160000.0 tokens per second
Current status
stated

The claim, verbatim

CoreWeave achieved 1.16 million tokens per second in server scenario on a single NVIDIA GB300 NVL72 rack for GPT-OSS-120B

Source (primary)

MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave (-, news_article)
View cached copy (2026-09-19)Live source ↗

Quote: “CoreWeave achieved over 1.16 million tokens per second in server”

How we checked this

Checked on September 24, 2026. The cited source supports every part of this claim.

The post states CoreWeave achieved over 1.16 million tokens per second in the server scenario on a single GB300 NVL72 rack for GPT-OSS-120B.

Confirmed in the source:

What we did: Read our cached copy of the publisher (https://www.coreweave.com/blog/coreweave-leads-cloud-providers-in-mlperf-r-inference-v6-1-performance-with-nvidia-blackwell-ultra) in full (14,189 characters, retrieved September 19, 2026) and checked each assertion in the claim against it.

Additional evidence

confirms MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave

Quote: “CoreWeave achieved over 1.16 million tokens per second in server”

View cached copy (2026-09-19)Live source ↗