jain.com

Full-rack DeepSeek R1 server throughput first place

Company: Nebius

Subject kind
product
Statement date
2026-09-16
Promised amount
603023.0 tokens/s
Current status
stated

The claim, verbatim

Nebius ranked first in server scenario for DeepSeek R1 on full-rack GB300 NVL72 system at 603,023 tokens per second

Source (primary)

MLPerf® Inference v6.1 Results: NVIDIA Vera Rubin NVL72 - Nebius (-, news_article)
View cached copy (2026-09-21)Live source ↗

Quote: “ranked first in both the server and offline scenarios for DeepSeek R1, at 603,023 and 689,961 tokens/s respectively”

How we checked this

Checked on September 25, 2026. The cited source supports every part of this claim.

Both the post's text and its table show a first-place DeepSeek R1 server result of 603,023 tokens/s on the 72-GPU GB300 NVL72.

Confirmed in the source:

What we did: Read our cached copy of the publisher (https://nebius.com/blog/posts/mlperf-inference-v6-1-results) in full (19,885 characters, retrieved September 21, 2026) and checked each assertion in the claim against it.

Additional evidence

confirms MLPerf® Inference v6.1 Results: NVIDIA Vera Rubin NVL72 - Nebius

Quote: “ranked first in both the server and offline scenarios for DeepSeek R1, at 603,023 and 689,961 tokens/s respectively”

View cached copy (2026-09-21)Live source ↗