Vera Rubin inference throughput gain vs GB200 (Cognition benchmark)
Company: CoreWeave
The claim, verbatim
Cognition measured up to a 4.8X increase in total token throughput for SWE-2 inference on Vera Rubin NVL72 versus a GB200 NVL72 baseline, in what CoreWeave described as the first customer-executed inference benchmark on the platform.
Source (primary)
CoreWeave brings Nvidia Vera Rubin, AI tools to its cloud services - Network World (-, news_article)
View cached copy (2026-10-02)Live source ↗Archive.org ↗
Quote: “it measured up to a 4.8X increase in total token throughput for SWE-2 inference workloads on Vera Rubin NVL72 compared with an Nvidia GB200 NVL72 baseline”
How we checked this
This claim has not yet been checked assertion-by-assertion against its source. It carries a cited source and quote, but the deeper check has not run. When it does, the result appears here whatever it says.
Additional evidence
confirms CoreWeave brings Nvidia Vera Rubin, AI tools to its cloud services - Network World
Quote: “it measured up to a 4.8X increase in total token throughput for SWE-2 inference workloads on Vera Rubin NVL72 compared with an Nvidia GB200 NVL72 baseline”
