Qwen3-VL-235B-A22B server throughput
Company: CoreWeave
The claim, verbatim
CoreWeave sustained 1,196 queries per second in server scenario with NVIDIA GB300 NVL72, the highest server throughput among cloud providers
Source (primary)
MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave (-, news_article)
View cached copy (2026-09-19)Live source ↗
Quote: “CoreWeave sustained 1,196 queries per second, the highest server throughput among cloud providers”
How we checked this
Checked on September 24, 2026. The cited source supports every part of this claim.
The post states CoreWeave sustained 1,196 queries per second in the Qwen3-VL-235B-A22B server scenario on an NVIDIA GB300 NVL72 rack. It says this was the highest server throughput among cloud providers.
Confirmed in the source:
- CoreWeave sustained 1,196 queries per second in the server scenario on Qwen3-VL-235B-A22B
- The result was achieved using an NVIDIA GB300 NVL72 rack
- This was the highest server throughput among cloud providers
What we did: Read our cached copy of the publisher (https://www.coreweave.com/blog/coreweave-leads-cloud-providers-in-mlperf-r-inference-v6-1-performance-with-nvidia-blackwell-ultra) in full (14,189 characters, retrieved September 19, 2026) and checked each assertion in the claim against it.
Additional evidence
confirms MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave
Quote: “CoreWeave sustained 1,196 queries per second, the highest server throughput among cloud providers”
