Full-rack gpt-oss 120B throughput
Company: Nebius
The claim, verbatim
Nebius full-rack system with NVIDIA GB300 NVL72 sustained over 1.1 million tokens per second on gpt-oss 120B
Source (primary)
MLPerf® Inference v6.1 Results: NVIDIA Vera Rubin NVL72 - Nebius (-, news_article)
View cached copy (2026-09-21)Live source ↗
Quote: “The same full-rack system sustained over 1.1 million tokens/s on gpt-oss 120B in both scenarios”
How we checked this
Checked on September 25, 2026. The cited source supports every part of this claim.
The post reports over 1.1 million tokens/s on gpt-oss 120B in both scenarios: 1,122,490 for server and 1,186,750 for offline.
Confirmed in the source:
- The full-rack GB300 NVL72 system sustained over 1.1 million tokens/s on gpt-oss 120B
What we did: Read our cached copy of the publisher (https://nebius.com/blog/posts/mlperf-inference-v6-1-results) in full (19,885 characters, retrieved September 21, 2026) and checked each assertion in the claim against it.
Additional evidence
confirms MLPerf® Inference v6.1 Results: NVIDIA Vera Rubin NVL72 - Nebius
Quote: “The same full-rack system sustained over 1.1 million tokens/s on gpt-oss 120B in both scenarios”
