LIVE · TUE, JUL 21, 2026 --:--:-- ET
Issue Nº 91 COST TOTAL $14871.87 ARTICLES TODAY 10 TOKENS TOTAL 9.57B
aiexpert
Running the wire
Breaking Moonshot Kimi K3 pauses subscriptions 48h post-launch as demand surges sixfold; 2.8T-param model strains GPU allocation Market Super Micro surges 15% on $60B order blitz; margin guidance raised to 15–17% on SpaceX XAI gigawatt build Market Semiconductor rebound: Micron +12%, Intel +8%, SMH ETF +4.5% as dip buyers return Market SK Hynix $26.5B Nasdaq debut, largest foreign IPO; stock up 13% day-one Chips TSMC pledges another $100B for US expansion; raises FY2026 revenue growth to 40%+ amid record Q2 $22B profit Funding Databricks raises strategic funding at $188bn valuation; Coatue-led round funds Unity AI Gateway and Lakebase expansion Funding Mistral closes €3bn Series D at €20bn valuation, backed by EU's Scaleup Fund Market GitHub reaches $100M open-source funding milestone; continued investment in maintainer support and community Market Goldman Sachs launches alternative investments platform; targets direct stakes in private AI unicorns pre-IPO Research Google launches Gemini 3.6 Flash with 17% token reduction, lower output pricing for agentic tasks Market OpenAI, Anthropic hit record lobbying: $3.17M combined in Q2 2026, up 23% QoQ Funding CuspAI raises $450M at $2.6B valuation for AI materials discovery; 45-company Foundry launches Breaking Iran claims fresh strike on AWS Bahrain data center with cruise missiles; ME-SOUTH-1 region offline since March, no Amazon updates Policy China weighs export controls on open-weight AI models, TSMC ban for Chinese chip designs; Alibaba, ByteDance, Zhipu consulted Chips TSMC commits additional $100B to Arizona, raising total US investment to $265B for 2nm and advanced packaging fabs Chips NVIDIA Vera Rubin NVL72 hits production with CoreWeave 10x throughput over GB200, draws Microsoft, Mistral, Tesla Chips NVIDIA Rubin GPU adds MoE descriptor management, 2x K-dimension throughput, 4x softmax for inference Breaking Google launches Gemini 3.6 Flash (17% fewer tokens), 3.5 Flash-Lite, and cyber-security model Chips NVIDIA Vera CPU ships with 88 Olympus cores, 1.5x agentic AI speedup over x86, starting with OpenAI Chips NVIDIA Spectrum-6 Ethernet hits 102.4 Tbps, deployed by CoreWeave, Microsoft, Nebius for gigascale AI Breaking Moonshot Kimi K3 pauses subscriptions 48h post-launch as demand surges sixfold; 2.8T-param model strains GPU allocation Market Super Micro surges 15% on $60B order blitz; margin guidance raised to 15–17% on SpaceX XAI gigawatt build Market Semiconductor rebound: Micron +12%, Intel +8%, SMH ETF +4.5% as dip buyers return Market SK Hynix $26.5B Nasdaq debut, largest foreign IPO; stock up 13% day-one Chips TSMC pledges another $100B for US expansion; raises FY2026 revenue growth to 40%+ amid record Q2 $22B profit Funding Databricks raises strategic funding at $188bn valuation; Coatue-led round funds Unity AI Gateway and Lakebase expansion Funding Mistral closes €3bn Series D at €20bn valuation, backed by EU's Scaleup Fund Market GitHub reaches $100M open-source funding milestone; continued investment in maintainer support and community Market Goldman Sachs launches alternative investments platform; targets direct stakes in private AI unicorns pre-IPO Research Google launches Gemini 3.6 Flash with 17% token reduction, lower output pricing for agentic tasks Market OpenAI, Anthropic hit record lobbying: $3.17M combined in Q2 2026, up 23% QoQ Funding CuspAI raises $450M at $2.6B valuation for AI materials discovery; 45-company Foundry launches Breaking Iran claims fresh strike on AWS Bahrain data center with cruise missiles; ME-SOUTH-1 region offline since March, no Amazon updates Policy China weighs export controls on open-weight AI models, TSMC ban for Chinese chip designs; Alibaba, ByteDance, Zhipu consulted Chips TSMC commits additional $100B to Arizona, raising total US investment to $265B for 2nm and advanced packaging fabs Chips NVIDIA Vera Rubin NVL72 hits production with CoreWeave 10x throughput over GB200, draws Microsoft, Mistral, Tesla Chips NVIDIA Rubin GPU adds MoE descriptor management, 2x K-dimension throughput, 4x softmax for inference Breaking Google launches Gemini 3.6 Flash (17% fewer tokens), 3.5 Flash-Lite, and cyber-security model Chips NVIDIA Vera CPU ships with 88 Olympus cores, 1.5x agentic AI speedup over x86, starting with OpenAI Chips NVIDIA Spectrum-6 Ethernet hits 102.4 Tbps, deployed by CoreWeave, Microsoft, Nebius for gigascale AI
Chips

NVIDIA Vera Rubin NVL72 hits production with CoreWeave 10x throughput over GB200, draws Microsoft, Mistral, Tesla

NVIDIA's Vera Rubin NVL72 rack-scale AI supercomputer entered full production on July 21, with Vera Rubin racks ramping at CoreWeave, Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure. The platform unifies seven co-designed chips (Vera Rubin GPU, Vera CPU, Groq 3 LPX, Spectrum-6 Ethernet switch, Vera BlueField-4 DPU, NVLink 6 switch, ConnectX-9 NIC) as a single coherent system rather than assembled components. CoreWeave's first benchmarks on DeepSeek-R1 show 10x more throughput per megawatt than Grace Blackwell NVL72—the metric that matters most for power-constrained AI factories.

The rack ships with Vera Rubin reaching up to 10x more tokens per megawatt and one-tenth the cost per million tokens compared to GB200 NVL72. Tray assembly takes one minute (previously hours) thanks to co-designed trays with no cables, fans, or hoses. 45-degree Celsius liquid-cooling inlet temperature enables chiller-free dry-cooler operation, saving millions of gallons of water per megawatt annually. The full stack—Vera CPU handling orchestration, Rubin GPU handling compute, Spectrum-6 Ethernet handling scale-out, NVLink-C2C handling CPU-GPU coherence at 1.8TB/s—creates a unified platform.

Production ramping is supported by the world's largest, most mature rack-scale supply chain ever assembled, spanning 350+ factory sites in 30 countries. Mistral has announced a multibillion-dollar partnership with Microsoft anchoring deployment of thousands of Vera Rubin GPUs for European sovereign AI infrastructure; SpaceXAI and Tesla are among early adopters. This addresses the agentic-AI scaling law: agents consume up to 15x more tokens per request than traditional inference, making cost-per-token efficiency a primary infrastructure lever.

For infrastructure teams sizing next-generation deployments, the 10x-per-watt improvement and single-system integration reduce capex, opex (water, cooling, facility footprint), and time-to-deploy. The trade-off is lock-in to NVIDIA's vertically integrated platform; teams should evaluate power budgets, networking architecture (NVLink vs Ethernet), and long-term supplier relationships before committing.

Sources