jain.com

Vera Rubin DeepSeek R1 leading per-GPU throughput

Company: Nebius

Subject kind
product
Statement date
2026-09-16
Promised amount
16427.0 tokens/s per GPU
Current status
stated

The claim, verbatim

Nebius posted the leading server result at 16,427 tokens per second per GPU on DeepSeek R1 with Nebius Vera Rubin NVL72

Source (primary)

MLPerf® Inference v6.1 Results: NVIDIA Vera Rubin NVL72 - Nebius (-, news_article)
View cached copy (2026-09-21)Live source ↗

Quote: “Nebius posted the leading server result at 16,427 tokens/s per GPU”

How we checked this

Checked on September 25, 2026. The cited source supports every part of this claim.

The post reports a leading per-GPU DeepSeek R1 server result of 16,427 tokens/s on the Vera Rubin NVL72 preview system. This figure is consistent with 591,368 tokens/s across 36 GPUs.

Confirmed in the source:

What we did: Read our cached copy of the publisher (https://nebius.com/blog/posts/mlperf-inference-v6-1-results) in full (19,885 characters, retrieved September 21, 2026) and checked each assertion in the claim against it.

Additional evidence

confirms MLPerf® Inference v6.1 Results: NVIDIA Vera Rubin NVL72 - Nebius

Quote: “Nebius posted the leading server result at 16,427 tokens/s per GPU”

View cached copy (2026-09-21)Live source ↗