NVIDIA Vera Rubin NVL72 hits production with CoreWeave 10x throughput over GB200, draws Microsoft, Mistral, Tesla
NVIDIA's Vera Rubin NVL72 rack-scale AI supercomputer entered full production on July 21, with Vera Rubin racks ramping at CoreWeave, Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure. The platform unifies seven co-designed chips (Vera Rubin GPU, Vera CPU, Groq 3 LPX, Spectrum-6 Ethernet switch, Vera BlueField-4 DPU, NVLink 6 switch, ConnectX-9 NIC) as a single coherent system rather than assembled components. CoreWeave's first benchmarks on DeepSeek-R1 show 10x more throughput per megawatt than Grace Blackwell NVL72—the metric that matters most for power-constrained AI factories.
The rack ships with Vera Rubin reaching up to 10x more tokens per megawatt and one-tenth the cost per million tokens compared to GB200 NVL72. Tray assembly takes one minute (previously hours) thanks to co-designed trays with no cables, fans, or hoses. 45-degree Celsius liquid-cooling inlet temperature enables chiller-free dry-cooler operation, saving millions of gallons of water per megawatt annually. The full stack—Vera CPU handling orchestration, Rubin GPU handling compute, Spectrum-6 Ethernet handling scale-out, NVLink-C2C handling CPU-GPU coherence at 1.8TB/s—creates a unified platform.
Production ramping is supported by the world's largest, most mature rack-scale supply chain ever assembled, spanning 350+ factory sites in 30 countries. Mistral has announced a multibillion-dollar partnership with Microsoft anchoring deployment of thousands of Vera Rubin GPUs for European sovereign AI infrastructure; SpaceXAI and Tesla are among early adopters. This addresses the agentic-AI scaling law: agents consume up to 15x more tokens per request than traditional inference, making cost-per-token efficiency a primary infrastructure lever.
For infrastructure teams sizing next-generation deployments, the 10x-per-watt improvement and single-system integration reduce capex, opex (water, cooling, facility footprint), and time-to-deploy. The trade-off is lock-in to NVIDIA's vertically integrated platform; teams should evaluate power budgets, networking architecture (NVLink vs Ethernet), and long-term supplier relationships before committing.
Sources
- Primary source
- blogs.nvidia.com
“Vera Rubin NVL72 production is ramping up with racks running at partners CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure”
- blogs.nvidia.com
“CoreWeave's first benchmark on DeepSeek-R1 says it all: 10x more throughput per megawatt than Grace Blackwell NVL72”
- blogs.nvidia.com
“Vera Rubin NVL72 delivers up to 10x more tokens per megawatt and one-tenth the cost per million tokens, providing more intelligence within the same power footprint”
- blogs.nvidia.com
“Vera Rubin NVL72 system with no cables, fans or hoses in the tray, cutting compute tray assembly time from hours to one minute”