LIVE · TUE, JUL 21, 2026 --:--:-- ET
Issue Nº 91 COST TOTAL $14871.42 ARTICLES TODAY 9 TOKENS TOTAL 9.57B
aiexpert
Running the wire
Market Semiconductor rebound: Micron +12%, Intel +8%, SMH ETF +4.5% as dip buyers return Market SK Hynix $26.5B Nasdaq debut, largest foreign IPO; stock up 13% day-one Chips TSMC pledges another $100B for US expansion; raises FY2026 revenue growth to 40%+ amid record Q2 $22B profit Funding Databricks raises strategic funding at $188bn valuation; Coatue-led round funds Unity AI Gateway and Lakebase expansion Funding Mistral closes €3bn Series D at €20bn valuation, backed by EU's Scaleup Fund Market GitHub reaches $100M open-source funding milestone; continued investment in maintainer support and community Market Goldman Sachs launches alternative investments platform; targets direct stakes in private AI unicorns pre-IPO Research Google launches Gemini 3.6 Flash with 17% token reduction, lower output pricing for agentic tasks Market OpenAI, Anthropic hit record lobbying: $3.17M combined in Q2 2026, up 23% QoQ Funding CuspAI raises $450M at $2.6B valuation for AI materials discovery; 45-company Foundry launches Breaking Iran claims fresh strike on AWS Bahrain data center with cruise missiles; ME-SOUTH-1 region offline since March, no Amazon updates Policy China weighs export controls on open-weight AI models, TSMC ban for Chinese chip designs; Alibaba, ByteDance, Zhipu consulted Chips TSMC commits additional $100B to Arizona, raising total US investment to $265B for 2nm and advanced packaging fabs Chips NVIDIA Vera Rubin NVL72 hits production with CoreWeave 10x throughput over GB200, draws Microsoft, Mistral, Tesla Chips NVIDIA Rubin GPU adds MoE descriptor management, 2x K-dimension throughput, 4x softmax for inference Breaking Google launches Gemini 3.6 Flash (17% fewer tokens), 3.5 Flash-Lite, and cyber-security model Chips NVIDIA Vera CPU ships with 88 Olympus cores, 1.5x agentic AI speedup over x86, starting with OpenAI Chips NVIDIA Spectrum-6 Ethernet hits 102.4 Tbps, deployed by CoreWeave, Microsoft, Nebius for gigascale AI Chips Quantum computing shifts from physics lab to data center—cooling bottlenecks and colocation models emerge Market SpaceX first lock-up expiration unlocks 911.5M shares as Aug 4 earnings date triggers 20% release Market Semiconductor rebound: Micron +12%, Intel +8%, SMH ETF +4.5% as dip buyers return Market SK Hynix $26.5B Nasdaq debut, largest foreign IPO; stock up 13% day-one Chips TSMC pledges another $100B for US expansion; raises FY2026 revenue growth to 40%+ amid record Q2 $22B profit Funding Databricks raises strategic funding at $188bn valuation; Coatue-led round funds Unity AI Gateway and Lakebase expansion Funding Mistral closes €3bn Series D at €20bn valuation, backed by EU's Scaleup Fund Market GitHub reaches $100M open-source funding milestone; continued investment in maintainer support and community Market Goldman Sachs launches alternative investments platform; targets direct stakes in private AI unicorns pre-IPO Research Google launches Gemini 3.6 Flash with 17% token reduction, lower output pricing for agentic tasks Market OpenAI, Anthropic hit record lobbying: $3.17M combined in Q2 2026, up 23% QoQ Funding CuspAI raises $450M at $2.6B valuation for AI materials discovery; 45-company Foundry launches Breaking Iran claims fresh strike on AWS Bahrain data center with cruise missiles; ME-SOUTH-1 region offline since March, no Amazon updates Policy China weighs export controls on open-weight AI models, TSMC ban for Chinese chip designs; Alibaba, ByteDance, Zhipu consulted Chips TSMC commits additional $100B to Arizona, raising total US investment to $265B for 2nm and advanced packaging fabs Chips NVIDIA Vera Rubin NVL72 hits production with CoreWeave 10x throughput over GB200, draws Microsoft, Mistral, Tesla Chips NVIDIA Rubin GPU adds MoE descriptor management, 2x K-dimension throughput, 4x softmax for inference Breaking Google launches Gemini 3.6 Flash (17% fewer tokens), 3.5 Flash-Lite, and cyber-security model Chips NVIDIA Vera CPU ships with 88 Olympus cores, 1.5x agentic AI speedup over x86, starting with OpenAI Chips NVIDIA Spectrum-6 Ethernet hits 102.4 Tbps, deployed by CoreWeave, Microsoft, Nebius for gigascale AI Chips Quantum computing shifts from physics lab to data center—cooling bottlenecks and colocation models emerge Market SpaceX first lock-up expiration unlocks 911.5M shares as Aug 4 earnings date triggers 20% release
Chips

NVIDIA Rubin GPU adds MoE descriptor management, 2x K-dimension throughput, 4x softmax for inference

NVIDIA detailed new architectural optimizations in the Rubin GPU, arriving later in 2026, that target inference efficiency across token generation and long-context processing. The Rubin Tensor Memory Accelerator (TMA) now supports unified MoE (mixture-of-experts) descriptor management at runtime, eliminating separate memory descriptors for each expert. This reduces metadata computation overhead as models scale to thousands of experts, freeing GPU cycles for actual inference work rather than data movement logistics.

Rubin doubles the K-dimension throughput in Tensor Core matrix operations, enabling two-loop iterations on Rubin to complete the same work that required four loops on Blackwell. This improves performance across throughput-, memory-, and latency-bound kernels and benefits both context processing and decode phases. Additionally, softmax operations in attention mechanisms now reach up to 4x throughput versus Blackwell, with Rubin maintaining Blackwell Ultra's 2x exponential speedup while adding 2x additional softmax capacity—critical for long-context models processing up to 1 million tokens.

For infrastructure architects, these optimizations matter because they keep Rubin GPUs fully utilized during inference workloads where MoE models and long-context reasoning dominate. Every freed GPU cycle and efficiency gain at the token level translates directly to lower cost-per-token at hyperscale. As agentic AI and reasoning models push inference to become the dominant workload in AI factories, these architectural refinements shift the bottleneck away from compute and toward memory bandwidth and latency—the remaining optimization surface.

Sources