Full-rack DeepSeek R1 offline throughput first place
Company: Nebius
The claim, verbatim
Nebius ranked first in offline scenario for DeepSeek R1 on full-rack GB300 NVL72 system at 689,961 tokens per second
Source (primary)
MLPerf® Inference v6.1 Results: NVIDIA Vera Rubin NVL72 - Nebius (-, news_article)
View cached copy (2026-09-21)Live source ↗
Quote: “ranked first in both the server and offline scenarios for DeepSeek R1, at 603,023 and 689,961 tokens/s respectively”
How we checked this
Checked on September 25, 2026. The cited source supports every part of this claim.
Both the post's text and its table show a first-place DeepSeek R1 offline result of 689,961 tokens/s on the 72-GPU GB300 NVL72.
Confirmed in the source:
- Nebius ranked first in the DeepSeek R1 offline scenario
- The result was on the full-rack GB300 NVL72 system
- Throughput was 689,961 tokens/s
What we did: Read our cached copy of the publisher (https://nebius.com/blog/posts/mlperf-inference-v6-1-results) in full (19,885 characters, retrieved September 21, 2026) and checked each assertion in the claim against it.
Additional evidence
confirms MLPerf® Inference v6.1 Results: NVIDIA Vera Rubin NVL72 - Nebius
Quote: “ranked first in both the server and offline scenarios for DeepSeek R1, at 603,023 and 689,961 tokens/s respectively”
