jain.com

GPT-OSS-120B per-GPU offline throughput

Company: CoreWeave

Subject kind
product
Statement date
2026-09-16
Promised amount
16635.0 tokens per second per GPU
Current status
stated

The claim, verbatim

CoreWeave sustained 16,635 tokens per second per GPU in offline scenario for GPT-OSS-120B, achieving the highest per-GPU throughput of any MLPerf v6.1 Datacenter Closed submission

Source (primary)

MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave (-, news_article)
View cached copy (2026-09-19)Live source ↗

Quote: “CoreWeave sustained 16,635 tokens per second per GPU in offline scenario”

How we checked this

Checked on September 24, 2026. The cited source supports every part of this claim.

The post reports 16,635 tokens per second per GPU offline on GPT-OSS-120B and calls it the highest per-GPU throughput of any v6.1 Datacenter Closed submission on that model. It also notes that per-GPU throughput is a derived metric not verified by MLCommons.

Confirmed in the source:

What we did: Read our cached copy of the publisher (https://www.coreweave.com/blog/coreweave-leads-cloud-providers-in-mlperf-r-inference-v6-1-performance-with-nvidia-blackwell-ultra) in full (14,189 characters, retrieved September 19, 2026) and checked each assertion in the claim against it.

Additional evidence

confirms MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave

Quote: “CoreWeave sustained 16,635 tokens per second per GPU in offline scenario”

View cached copy (2026-09-19)Live source ↗