jain.com

Full-rack gpt-oss 120B throughput

Company: Nebius

Subject kind
product
Statement date
2026-09-16
Promised amount
1100000.0 tokens/s
Current status
stated

The claim, verbatim

Nebius full-rack system with NVIDIA GB300 NVL72 sustained over 1.1 million tokens per second on gpt-oss 120B

Source (primary)

MLPerf® Inference v6.1 Results: NVIDIA Vera Rubin NVL72 - Nebius (-, news_article)
View cached copy (2026-09-21)Live source ↗

Quote: “The same full-rack system sustained over 1.1 million tokens/s on gpt-oss 120B in both scenarios”

How we checked this

Checked on September 25, 2026. The cited source supports every part of this claim.

The post reports over 1.1 million tokens/s on gpt-oss 120B in both scenarios: 1,122,490 for server and 1,186,750 for offline.

Confirmed in the source:

What we did: Read our cached copy of the publisher (https://nebius.com/blog/posts/mlperf-inference-v6-1-results) in full (19,885 characters, retrieved September 21, 2026) and checked each assertion in the claim against it.

Additional evidence

confirms MLPerf® Inference v6.1 Results: NVIDIA Vera Rubin NVL72 - Nebius

Quote: “The same full-rack system sustained over 1.1 million tokens/s on gpt-oss 120B in both scenarios”

View cached copy (2026-09-21)Live source ↗