Vera Rubin NVL72 4.8x inference throughput vs GB200 for Cognition
Company: CoreWeave
The claim, verbatim
Cognition's benchmarks on CoreWeave showed up to 4.8x total token throughput per GPU for SWE-2 inference on Vera Rubin NVL72 versus GB200 NVL72 at matched interactivity.
Source (primary)
First Vera Rubin NVL72 Customer Sees 4.8x Throughput - CoreWeave (-, news_article)
View cached copy (2026-10-01)Live source ↗Archive.org ↗
Quote: “our engineers are seeing up to a 4.8x increase in total token throughput for SWE-2 inference workloads over GB200 NVL72”
How we checked this
This claim has not yet been checked assertion-by-assertion against its source. It carries a cited source and quote, but the deeper check has not run. When it does, the result appears here whatever it says.
Additional evidence
confirms First Vera Rubin NVL72 Customer Sees 4.8x Throughput - CoreWeave
Quote: “our engineers are seeing up to a 4.8x increase in total token throughput for SWE-2 inference workloads over GB200 NVL72”
