CoreWeave Powers First Vera Rubin Rack; Inference Cost Per Token Slashes 10x vs. Blackwell in Q3 Rollout
CoreWeave completed the first production Vera Rubin NVL72 system bring-up on June 1, 2026, marking the industry's first validated deployment of NVIDIA's next-generation rack-scale AI platform. The system is now running live compute jobs. Vera Rubin delivers 3.6 EFLOPS of NVFP4 inference performance per rack (5x Blackwell) and 10x lower cost per token, with claims ranging from 70–90% token-cost reduction for 70B+ parameter models at high concurrency. The platform combines 72 Rubin GPUs (336B transistors each, 50 PFLOPS FP4 per GPU) and 36 Vera CPUs with custom Olympus ARM cores for low-latency agentic workloads. H2 2026 availability is confirmed with AWS, Google Cloud, Microsoft Azure, and OCI.
The 10x per-token efficiency comes from codesigned hardware: 22 TB/s memory bandwidth per GPU (vs. 8 TB/s on Blackwell), NVLink 6 delivering 260 TB/s rack-scale bandwidth, and photonics-enabled Spectrum-X Ethernet. For workload-specific gains: 7B–13B models are memory-bandwidth-bound, yielding 2–3x throughput per dollar; 405B+ models running FP4 are compute-bound, achieving near-10x gains. Initial cloud pricing expected at 30–50% premium over Blackwell, but breakeven economics favor large-scale inference deployments. Rubin Ultra (144 GPUs, 15 EFLOPS) targets H2 2027.
For ops and infrastructure teams, this marks the inflection from training-bound to inference-bound economics. At CoreWeave's scale and with Microsoft/Mistral anchoring European capacity, Vera Rubin racks will be allocation-gated through 2027. Practical access for most teams lands in 2027; immediate-term focus should be Blackwell reserve capacity contracts or spot pricing. The premium pricing reflects both scarcity and the real efficiency gain—validate your token-per-megawatt assumptions against Rubin benchmarks before 2H2026 cutoff.