MLPerf v6.1 Llama 2 70B inference throughput leadership
Company: CoreWeave
The claim, verbatim
CoreWeave delivered the highest throughput among cloud providers for Llama 2 70B on NVIDIA GB300 NVL72, achieving 944,902 tokens per second in server scenario and 1,136,100 in offline scenario
Source (primary)
MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave (-, news_article)
View cached copy (2026-09-19)Live source ↗
Quote: “We delivered the highest throughput among cloud providers with NVIDIA GB300 NVL72 with a throughput of 944,902 tokens per second in server scenario and 1,136,100 in offline scenario.”
How we checked this
Checked on September 24, 2026. The cited source supports every part of this claim.
The blog states that CoreWeave had the highest Llama 2 70B throughput among cloud providers on GB300 NVL72, with 944,902 tokens per second in the server scenario and 1,136,100 in the offline scenario.
Confirmed in the source:
- CoreWeave delivered the highest Llama 2 70B throughput among cloud providers using NVIDIA GB300 NVL72
- Llama 2 70B throughput of 944,902 tokens per second in the server scenario
- Llama 2 70B throughput of 1,136,100 tokens per second in the offline scenario
What we did: Read our cached copy of the publisher (https://www.coreweave.com/blog/coreweave-leads-cloud-providers-in-mlperf-r-inference-v6-1-performance-with-nvidia-blackwell-ultra) in full (14,189 characters, retrieved September 19, 2026) and checked each assertion in the claim against it.
Additional evidence
confirms MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave
Quote: “We delivered the highest throughput among cloud providers with NVIDIA GB300 NVL72 with a throughput of 944,902 tokens per second in server scenario and 1,136,100 in offline scenario.”
