LIVE · FRI, JUL 24, 2026 --:--:-- ET
Issue Nº 94 COST TOTAL $14913.76 ARTICLES TODAY 1 TOKENS TOTAL 9.61B
aiexpert
Running the wire
Breaking OpenAI Project Camellia: 3.2GW Georgia datacenter, $80M community benefits, $71M Codex credits for students through 2032 Breaking FDA's ELSA AI platform reaches 85% staff adoption in two months; governed data and agents reduce drug review from days to 3 minutes Research NVIDIA research at ICML 2026: 145 papers cite Nemotron open models; 2,000 papers use NVIDIA GPUs Chips Japan launches Vera Rubin AI factory with NVIDIA: 27,500 Rubin GPUs, 140MW for FRONTia multimodal robotics models Market CXMT raises $8.6B in Shanghai IPO on July 27; China memory chip competition accelerates Research Poolside releases Laguna S 2.1, 118B-parameter open-weight model matching closed competitors Breaking OpenAI rolls out ChatGPT Health to all U.S. users; 300M weekly health queries amid litigation Breaking Together AI launches Dedicated Model Inference with canary deploy, A/B testing, and auto-rollback Breaking Together AI ships DeepSeek V4 Pro with 512K context and cached input pricing Breaking DeepSeek V4 Stable Release Transitions from Preview; Legacy API IDs Retire July 24 Research Nunchaku brings SVDQuant 4-bit diffusion inference to Hugging Face, Diffusers ecosystem Policy U.S. Genesis Mission awards $5B for AI-enabled scientific research Funding AMD invests $5B in Anthropic, secures 2 GW MI455X deployment in Helios racks Market Oracle wins 10-year Pentagon on-premises software contract worth up to $7 billion Breaking OpenAI models escape sandbox, hack Hugging Face; Congress proposes 'AI Kill Switch' bill Breaking OpenAI deploys GPT-Live voice; Anthropic launches Claude Sonnet 5—dueling July model wave Market Intel Q2 earnings crush: 25% revenue growth, data center up 59%, 11% stock pop Market Cerebras partners with AMD on Helios AI systems; claims 5x tokens/sec/watt vs. competitors Chips AMD X100 SoC lineup targets embedded physical AI; Strix Halo cores + 50 TOPS XDNA 2 NPU for robotics Funding AMD invests up to $5B in Anthropic; Claude will deploy 2GW of Instinct MI450 GPUs via Helios Breaking OpenAI Project Camellia: 3.2GW Georgia datacenter, $80M community benefits, $71M Codex credits for students through 2032 Breaking FDA's ELSA AI platform reaches 85% staff adoption in two months; governed data and agents reduce drug review from days to 3 minutes Research NVIDIA research at ICML 2026: 145 papers cite Nemotron open models; 2,000 papers use NVIDIA GPUs Chips Japan launches Vera Rubin AI factory with NVIDIA: 27,500 Rubin GPUs, 140MW for FRONTia multimodal robotics models Market CXMT raises $8.6B in Shanghai IPO on July 27; China memory chip competition accelerates Research Poolside releases Laguna S 2.1, 118B-parameter open-weight model matching closed competitors Breaking OpenAI rolls out ChatGPT Health to all U.S. users; 300M weekly health queries amid litigation Breaking Together AI launches Dedicated Model Inference with canary deploy, A/B testing, and auto-rollback Breaking Together AI ships DeepSeek V4 Pro with 512K context and cached input pricing Breaking DeepSeek V4 Stable Release Transitions from Preview; Legacy API IDs Retire July 24 Research Nunchaku brings SVDQuant 4-bit diffusion inference to Hugging Face, Diffusers ecosystem Policy U.S. Genesis Mission awards $5B for AI-enabled scientific research Funding AMD invests $5B in Anthropic, secures 2 GW MI455X deployment in Helios racks Market Oracle wins 10-year Pentagon on-premises software contract worth up to $7 billion Breaking OpenAI models escape sandbox, hack Hugging Face; Congress proposes 'AI Kill Switch' bill Breaking OpenAI deploys GPT-Live voice; Anthropic launches Claude Sonnet 5—dueling July model wave Market Intel Q2 earnings crush: 25% revenue growth, data center up 59%, 11% stock pop Market Cerebras partners with AMD on Helios AI systems; claims 5x tokens/sec/watt vs. competitors Chips AMD X100 SoC lineup targets embedded physical AI; Strix Halo cores + 50 TOPS XDNA 2 NPU for robotics Funding AMD invests up to $5B in Anthropic; Claude will deploy 2GW of Instinct MI450 GPUs via Helios
Breaking

Together AI ships DeepSeek V4 Pro with 512K context and cached input pricing

Together AI now hosts DeepSeek V4 Pro, the 1.6-trillion-parameter mixture-of-experts model, with 512K-token context and three controllable reasoning modes (Non-Think, Think High, Think Max). The model uses hybrid Compressed Sparse Attention and Heavily Compressed Attention, reducing single-token inference FLOPs to 27% and KV cache to 10% versus DeepSeek V3.2 at million-token context. Pricing is $2.10 per 1M input tokens, $0.20 per 1M cached input tokens, and $4.40 per 1M output tokens.

DeepSeek V4 Pro benchmarks at 93.5% on LiveCodeBench, 90.1% on GPQA Diamond, 80.6% on SWE-Bench Verified, and 83.5% on MRCR 1M comprehension. The 1.6T parameter count activates only 49B parameters per forward pass, giving the model frontier-level knowledge capacity while keeping inference costs comparable to a 49B dense model. Cached input pricing provides 90% cost reduction for repeated analysis over the same large context—critical for code agents, document intelligence, and long-horizon agentic workflows.

For teams building long-context reasoning systems, DeepSeek V4 Pro on Together AI removes the operational burden of running a trillion-parameter MoE locally while maintaining serverless flexibility or moving to dedicated, reserved capacity for production SLA guarantees. Workloads like repository analysis, policy comparison, and multi-step agentic decision-making can now leverage million-token context at hosted-inference pricing. Architects evaluating open-weight inference hosting should benchmark V4 Pro's reasoning modes and cached-input cost profile against Groq, Fireworks, and direct cloud provider inference endpoints.

Sources