Llama 2 70B on GB300 NVL72 offline
Company: CoreWeave
The claim, verbatim
CoreWeave achieved 1,136,100 tokens per second in offline scenario with NVIDIA GB300 NVL72 for Llama 2 70B, the only cloud provider result above one million tokens per second in this scenario
Source (primary)
MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave (-, news_article)
View cached copy (2026-09-19)Live source ↗
Quote: “1,136,100 in offline scenario. In the offline scenario, we achieved the only result above one million tokens per second among cloud provider submissions using NVIDIA GB300 NVL72.”
How we checked this
Checked on September 24, 2026. The cited source supports part of this claim, but not all of it.
The post confirms the 1,136,100 tokens-per-second offline figure. However, it limits the 'only result above one million' claim to cloud provider submissions using GB300 NVL72, not all cloud provider results in the scenario.
Confirmed in the source:
- CoreWeave achieved 1,136,100 tokens per second in the offline scenario on Llama 2 70B with NVIDIA GB300 NVL72
- It was the only result above one million tokens per second among cloud provider submissions using NVIDIA GB300 NVL72
We could not confirm this from the cited source:
- That it was the only cloud provider result above one million tokens per second in the Llama 2 70B offline scenario across all hardware platforms
That does not mean it is false — only that this source does not establish it, as of the date above. If we find a source that settles it, we will re-check and update this page.
What we did: Read our cached copy of the publisher (https://www.coreweave.com/blog/coreweave-leads-cloud-providers-in-mlperf-r-inference-v6-1-performance-with-nvidia-blackwell-ultra) in full (14,189 characters, retrieved September 19, 2026) and checked each assertion in the claim against it.
Additional evidence
confirms MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave
Quote: “1,136,100 in offline scenario. In the offline scenario, we achieved the only result above one million tokens per second among cloud provider submissions using NVIDIA GB300 NVL72.”
