Vera Rubin DeepSeek R1 leading per-GPU throughput
Company: Nebius
The claim, verbatim
Nebius posted the leading server result at 16,427 tokens per second per GPU on DeepSeek R1 with Nebius Vera Rubin NVL72
Source (primary)
MLPerf® Inference v6.1 Results: NVIDIA Vera Rubin NVL72 - Nebius (-, news_article)
View cached copy (2026-09-21)Live source ↗
Quote: “Nebius posted the leading server result at 16,427 tokens/s per GPU”
How we checked this
Checked on September 25, 2026. The cited source supports every part of this claim.
The post reports a leading per-GPU DeepSeek R1 server result of 16,427 tokens/s on the Vera Rubin NVL72 preview system. This figure is consistent with 591,368 tokens/s across 36 GPUs.
Confirmed in the source:
- Nebius posted the leading server result at 16,427 tokens/s per GPU
- The result was on DeepSeek R1
- The result was on the Nebius Vera Rubin NVL72 system
What we did: Read our cached copy of the publisher (https://nebius.com/blog/posts/mlperf-inference-v6-1-results) in full (19,885 characters, retrieved September 21, 2026) and checked each assertion in the claim against it.
Additional evidence
confirms MLPerf® Inference v6.1 Results: NVIDIA Vera Rubin NVL72 - Nebius
Quote: “Nebius posted the leading server result at 16,427 tokens/s per GPU”
