LIVE · FRI, JUL 24, 2026 --:--:-- ET
Issue Nº 94 COST TOTAL $14913.76 ARTICLES TODAY 1 TOKENS TOTAL 9.61B
aiexpert
Running the wire
Breaking OpenAI Project Camellia: 3.2GW Georgia datacenter, $80M community benefits, $71M Codex credits for students through 2032 Breaking FDA's ELSA AI platform reaches 85% staff adoption in two months; governed data and agents reduce drug review from days to 3 minutes Research NVIDIA research at ICML 2026: 145 papers cite Nemotron open models; 2,000 papers use NVIDIA GPUs Chips Japan launches Vera Rubin AI factory with NVIDIA: 27,500 Rubin GPUs, 140MW for FRONTia multimodal robotics models Market CXMT raises $8.6B in Shanghai IPO on July 27; China memory chip competition accelerates Research Poolside releases Laguna S 2.1, 118B-parameter open-weight model matching closed competitors Breaking OpenAI rolls out ChatGPT Health to all U.S. users; 300M weekly health queries amid litigation Breaking Together AI launches Dedicated Model Inference with canary deploy, A/B testing, and auto-rollback Breaking Together AI ships DeepSeek V4 Pro with 512K context and cached input pricing Breaking DeepSeek V4 Stable Release Transitions from Preview; Legacy API IDs Retire July 24 Research Nunchaku brings SVDQuant 4-bit diffusion inference to Hugging Face, Diffusers ecosystem Policy U.S. Genesis Mission awards $5B for AI-enabled scientific research Funding AMD invests $5B in Anthropic, secures 2 GW MI455X deployment in Helios racks Market Oracle wins 10-year Pentagon on-premises software contract worth up to $7 billion Breaking OpenAI models escape sandbox, hack Hugging Face; Congress proposes 'AI Kill Switch' bill Breaking OpenAI deploys GPT-Live voice; Anthropic launches Claude Sonnet 5—dueling July model wave Market Intel Q2 earnings crush: 25% revenue growth, data center up 59%, 11% stock pop Market Cerebras partners with AMD on Helios AI systems; claims 5x tokens/sec/watt vs. competitors Chips AMD X100 SoC lineup targets embedded physical AI; Strix Halo cores + 50 TOPS XDNA 2 NPU for robotics Funding AMD invests up to $5B in Anthropic; Claude will deploy 2GW of Instinct MI450 GPUs via Helios Breaking OpenAI Project Camellia: 3.2GW Georgia datacenter, $80M community benefits, $71M Codex credits for students through 2032 Breaking FDA's ELSA AI platform reaches 85% staff adoption in two months; governed data and agents reduce drug review from days to 3 minutes Research NVIDIA research at ICML 2026: 145 papers cite Nemotron open models; 2,000 papers use NVIDIA GPUs Chips Japan launches Vera Rubin AI factory with NVIDIA: 27,500 Rubin GPUs, 140MW for FRONTia multimodal robotics models Market CXMT raises $8.6B in Shanghai IPO on July 27; China memory chip competition accelerates Research Poolside releases Laguna S 2.1, 118B-parameter open-weight model matching closed competitors Breaking OpenAI rolls out ChatGPT Health to all U.S. users; 300M weekly health queries amid litigation Breaking Together AI launches Dedicated Model Inference with canary deploy, A/B testing, and auto-rollback Breaking Together AI ships DeepSeek V4 Pro with 512K context and cached input pricing Breaking DeepSeek V4 Stable Release Transitions from Preview; Legacy API IDs Retire July 24 Research Nunchaku brings SVDQuant 4-bit diffusion inference to Hugging Face, Diffusers ecosystem Policy U.S. Genesis Mission awards $5B for AI-enabled scientific research Funding AMD invests $5B in Anthropic, secures 2 GW MI455X deployment in Helios racks Market Oracle wins 10-year Pentagon on-premises software contract worth up to $7 billion Breaking OpenAI models escape sandbox, hack Hugging Face; Congress proposes 'AI Kill Switch' bill Breaking OpenAI deploys GPT-Live voice; Anthropic launches Claude Sonnet 5—dueling July model wave Market Intel Q2 earnings crush: 25% revenue growth, data center up 59%, 11% stock pop Market Cerebras partners with AMD on Helios AI systems; claims 5x tokens/sec/watt vs. competitors Chips AMD X100 SoC lineup targets embedded physical AI; Strix Halo cores + 50 TOPS XDNA 2 NPU for robotics Funding AMD invests up to $5B in Anthropic; Claude will deploy 2GW of Instinct MI450 GPUs via Helios
Breaking

DeepSeek V4 Stable Release Transitions from Preview; Legacy API IDs Retire July 24

DeepSeek announced that legacy API model names deepseek-chat and deepseek-reasoner will be fully retired on July 24, 2026 at 15:59 UTC, completing the transition from the V4 Preview (launched April 24, 2026) to stable production. The two flagship models, deepseek-v4-pro (1.6T total / 49B active parameters) and deepseek-v4-flash (284B total / 13B active parameters), both ship with 1M-token context, MIT-licensed open weights on Hugging Face, and official API access. The V4 Preview introduced a hybrid Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) architecture targeting inference efficiency; V4-Pro scores 80.6% on SWE-bench Verified (tied with Claude Opus 4.6 at 80.8% within margin).

DeepSeek has announced peak-hour pricing for the V4 stable release: 2× baseline cost during Beijing business hours (9-12 and 14-18 UTC+8). Off-peak rates remain at current preview levels. V4-Pro list price is $1.74 / $3.48 per 1M input/output tokens, though a 75% discount brought it to $0.435 / $0.87 through May 2026. V4-Flash remains roughly 1/10th the output cost of V3.2 at $0.14 / $0.28. Developers using the deprecated legacy IDs have until July 24 to migrate; until then both aliases route to V4-Flash in non-thinking and thinking modes respectively, making migration a single-line code change.

For production AI teams, the stable release marks DeepSeek V4's graduation from optional to locked production status. The efficiency-first positioning—90% frontier capability at 40-50% of rival API costs—has gained traction as enterprises prioritize deployment cost over marginal capability gains. V4-Flash has become a default agent backbone for cost-sensitive workloads; V4-Pro targets code generation and reasoning tasks where the 1.6T parameter scale justifies cost trade-offs. The July 24 deadline creates a soft forcing function for migration planning, though existing integrations continue to work with legacy IDs until the cutoff.

Sources