jain.com

MLPerf v6.1 Qwen3-VL multimodal inference throughput

Company: CoreWeave

Subject kind
product
Statement date
2026-09-16
Current status
stated

The claim, verbatim

CoreWeave achieved the highest server throughput among cloud providers for Qwen3-VL-235B multimodal model on NVIDIA GB300 NVL72, sustaining 1,196 queries per second in server scenario and 1,135 samples per second in offline scenario

Source (primary)

MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave (-, news_article)
View cached copy (2026-09-19)Live source ↗

Quote: “CoreWeave delivered the highest server throughput among cloud providers using an NVIDIA GB300 NVL72 rack. In server scenario, CoreWeave sustained 1,196 queries per second, the highest server throughput among cloud providers, and 1,135 samples per second in offline scenario.”

How we checked this

Checked on September 24, 2026. The cited source supports every part of this claim.

The post states CoreWeave delivered the highest server throughput among cloud providers on the multimodal Qwen3-VL-235B-A22B using a GB300 NVL72 rack. It reports 1,196 queries per second in the server scenario and 1,135 samples per second offline.

Confirmed in the source:

What we did: Read our cached copy of the publisher (https://www.coreweave.com/blog/coreweave-leads-cloud-providers-in-mlperf-r-inference-v6-1-performance-with-nvidia-blackwell-ultra) in full (14,189 characters, retrieved September 19, 2026) and checked each assertion in the claim against it.

Additional evidence

confirms MLPerf® Inference v6.1 Results: CoreWeave Leads Providers - CoreWeave

Quote: “CoreWeave delivered the highest server throughput among cloud providers using an NVIDIA GB300 NVL72 rack. In server scenario, CoreWeave sustained 1,196 queries per second, the highest server throughput among cloud providers, and 1,135 samples per second in offline scenario.”

View cached copy (2026-09-19)Live source ↗