aiexpert
Home / News / Brief
Chips · Jul 21, 2026, 04:03 PM · 5 sources

NVIDIA Vera Rubin NVL72 hits production with CoreWeave 10x throughput over GB200, draws Microsoft, Mistral, Tesla

NVIDIA's Vera Rubin NVL72 rack-scale AI supercomputer entered full production on July 21, with Vera Rubin racks ramping at CoreWeave, Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure. The platform unifies seven co-designed chips (Vera Rubin GPU, Vera CPU, Groq 3 LPX, Spectrum-6 Ethernet switch, Vera BlueField-4 DPU, NVLink 6 switch, ConnectX-9 NIC) as a single coherent system rather than assembled components. CoreWeave's first benchmarks on DeepSeek-R1 show 10x more throughput per megawatt than Grace Blackwell NVL72—the metric that matters most for power-constrained AI factories.

The rack ships with Vera Rubin reaching up to 10x more tokens per megawatt and one-tenth the cost per million tokens compared to GB200 NVL72. Tray assembly takes one minute (previously hours) thanks to co-designed trays with no cables, fans, or hoses. 45-degree Celsius liquid-cooling inlet temperature enables chiller-free dry-cooler operation, saving millions of gallons of water per megawatt annually. The full stack—Vera CPU handling orchestration, Rubin GPU handling compute, Spectrum-6 Ethernet handling scale-out, NVLink-C2C handling CPU-GPU coherence at 1.8TB/s—creates a unified platform.

Production ramping is supported by the world's largest, most mature rack-scale supply chain ever assembled, spanning 350+ factory sites in 30 countries. Mistral has announced a multibillion-dollar partnership with Microsoft anchoring deployment of thousands of Vera Rubin GPUs for European sovereign AI infrastructure; SpaceXAI and Tesla are among early adopters. This addresses the agentic-AI scaling law: agents consume up to 15x more tokens per request than traditional inference, making cost-per-token efficiency a primary infrastructure lever.

For infrastructure teams sizing next-generation deployments, the 10x-per-watt improvement and single-system integration reduce capex, opex (water, cooling, facility footprint), and time-to-deploy. The trade-off is lock-in to NVIDIA's vertically integrated platform; teams should evaluate power budgets, networking architecture (NVLink vs Ethernet), and long-term supplier relationships before committing.

Sources

Everything this brief rests on
  1. 01 Primary source blogs.nvidia.com
  2. 02 blogs.nvidia.com blogs.nvidia.com “Vera Rubin NVL72 production is ramping up with racks running at partners CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure”
  3. 03 blogs.nvidia.com blogs.nvidia.com “CoreWeave's first benchmark on DeepSeek-R1 says it all: 10x more throughput per megawatt than Grace Blackwell NVL72”
  4. 04 blogs.nvidia.com blogs.nvidia.com “Vera Rubin NVL72 delivers up to 10x more tokens per megawatt and one-tenth the cost per million tokens, providing more intelligence within the same power footprint”
  5. 05 blogs.nvidia.com blogs.nvidia.com “Vera Rubin NVL72 system with no cables, fans or hoses in the tray, cutting compute tray assembly time from hours to one minute”