Llama 2 70B on GB300 NVL72 server
Company: CoreWeave
The claim, verbatim
CoreWeave delivered 944,902 tokens per second in server scenario with NVIDIA GB300 NVL72 for Llama 2 70B, the highest throughput among cloud providers
Source (primary)
MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave (-, news_article)
View cached copy (2026-09-19)Live source ↗
Quote: “We delivered the highest throughput among cloud providers with NVIDIA GB300 NVL72 with a throughput of 944,902 tokens per second in server scenario”
How we checked this
Checked on September 24, 2026. The cited source supports every part of this claim.
The post states CoreWeave delivered the highest throughput among cloud providers with GB300 NVL72 on Llama 2 70B, at 944,902 tokens per second in the server scenario.
Confirmed in the source:
- CoreWeave delivered 944,902 tokens per second in the server scenario on Llama 2 70B
- The result used NVIDIA GB300 NVL72
- This was the highest throughput among cloud providers with NVIDIA GB300 NVL72
What we did: Read our cached copy of the publisher (https://www.coreweave.com/blog/coreweave-leads-cloud-providers-in-mlperf-r-inference-v6-1-performance-with-nvidia-blackwell-ultra) in full (14,189 characters, retrieved September 19, 2026) and checked each assertion in the claim against it.
Additional evidence
confirms MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave
Quote: “We delivered the highest throughput among cloud providers with NVIDIA GB300 NVL72 with a throughput of 944,902 tokens per second in server scenario”
