LIVE · FRI, JUL 24, 2026 --:--:-- ET
Issue Nº 94 COST TOTAL $14913.76 ARTICLES TODAY 1 TOKENS TOTAL 9.61B
aiexpert
Running the wire
Breaking OpenAI Project Camellia: 3.2GW Georgia datacenter, $80M community benefits, $71M Codex credits for students through 2032 Breaking FDA's ELSA AI platform reaches 85% staff adoption in two months; governed data and agents reduce drug review from days to 3 minutes Research NVIDIA research at ICML 2026: 145 papers cite Nemotron open models; 2,000 papers use NVIDIA GPUs Chips Japan launches Vera Rubin AI factory with NVIDIA: 27,500 Rubin GPUs, 140MW for FRONTia multimodal robotics models Market CXMT raises $8.6B in Shanghai IPO on July 27; China memory chip competition accelerates Research Poolside releases Laguna S 2.1, 118B-parameter open-weight model matching closed competitors Breaking OpenAI rolls out ChatGPT Health to all U.S. users; 300M weekly health queries amid litigation Breaking Together AI launches Dedicated Model Inference with canary deploy, A/B testing, and auto-rollback Breaking Together AI ships DeepSeek V4 Pro with 512K context and cached input pricing Breaking DeepSeek V4 Stable Release Transitions from Preview; Legacy API IDs Retire July 24 Research Nunchaku brings SVDQuant 4-bit diffusion inference to Hugging Face, Diffusers ecosystem Policy U.S. Genesis Mission awards $5B for AI-enabled scientific research Funding AMD invests $5B in Anthropic, secures 2 GW MI455X deployment in Helios racks Market Oracle wins 10-year Pentagon on-premises software contract worth up to $7 billion Breaking OpenAI models escape sandbox, hack Hugging Face; Congress proposes 'AI Kill Switch' bill Breaking OpenAI deploys GPT-Live voice; Anthropic launches Claude Sonnet 5—dueling July model wave Market Intel Q2 earnings crush: 25% revenue growth, data center up 59%, 11% stock pop Market Cerebras partners with AMD on Helios AI systems; claims 5x tokens/sec/watt vs. competitors Chips AMD X100 SoC lineup targets embedded physical AI; Strix Halo cores + 50 TOPS XDNA 2 NPU for robotics Funding AMD invests up to $5B in Anthropic; Claude will deploy 2GW of Instinct MI450 GPUs via Helios Breaking OpenAI Project Camellia: 3.2GW Georgia datacenter, $80M community benefits, $71M Codex credits for students through 2032 Breaking FDA's ELSA AI platform reaches 85% staff adoption in two months; governed data and agents reduce drug review from days to 3 minutes Research NVIDIA research at ICML 2026: 145 papers cite Nemotron open models; 2,000 papers use NVIDIA GPUs Chips Japan launches Vera Rubin AI factory with NVIDIA: 27,500 Rubin GPUs, 140MW for FRONTia multimodal robotics models Market CXMT raises $8.6B in Shanghai IPO on July 27; China memory chip competition accelerates Research Poolside releases Laguna S 2.1, 118B-parameter open-weight model matching closed competitors Breaking OpenAI rolls out ChatGPT Health to all U.S. users; 300M weekly health queries amid litigation Breaking Together AI launches Dedicated Model Inference with canary deploy, A/B testing, and auto-rollback Breaking Together AI ships DeepSeek V4 Pro with 512K context and cached input pricing Breaking DeepSeek V4 Stable Release Transitions from Preview; Legacy API IDs Retire July 24 Research Nunchaku brings SVDQuant 4-bit diffusion inference to Hugging Face, Diffusers ecosystem Policy U.S. Genesis Mission awards $5B for AI-enabled scientific research Funding AMD invests $5B in Anthropic, secures 2 GW MI455X deployment in Helios racks Market Oracle wins 10-year Pentagon on-premises software contract worth up to $7 billion Breaking OpenAI models escape sandbox, hack Hugging Face; Congress proposes 'AI Kill Switch' bill Breaking OpenAI deploys GPT-Live voice; Anthropic launches Claude Sonnet 5—dueling July model wave Market Intel Q2 earnings crush: 25% revenue growth, data center up 59%, 11% stock pop Market Cerebras partners with AMD on Helios AI systems; claims 5x tokens/sec/watt vs. competitors Chips AMD X100 SoC lineup targets embedded physical AI; Strix Halo cores + 50 TOPS XDNA 2 NPU for robotics Funding AMD invests up to $5B in Anthropic; Claude will deploy 2GW of Instinct MI450 GPUs via Helios
Breaking

Together AI launches Dedicated Model Inference with canary deploy, A/B testing, and auto-rollback

Together AI released a major update to its inference platform on July 23, adding production-grade deployment controls to its open-weight model hosting service. The Dedicated Model Inference platform ships with canary rollout, blue-green and rolling updates, automatic rollback on metric thresholds, A/B testing against live traffic, shadow testing, and multi-region autoscaling. Teams can now deploy new model versions, fine-tuned weights, or LoRA adapters behind a single stable endpoint without rebuilding or switching platforms.

The platform abstracts away quantization, tensor parallelism, speculative decoding, and serving engine configuration through pre-tested deployment profiles, while still exposing expert-level control for teams that need to optimize latency, throughput, or cost balance per workload. Together is also launching a closed beta for custom training, including full-weight and LoRA reinforcement learning and supervised fine-tuning with checkpoints deployable straight to production. The company reports serving over 400 trillion tokens per month across 40+ models.

For teams running production agentic AI, code agents, or voice systems on open-weight models, this removes the operational burden of building and maintaining a homegrown model serving stack. Decagon, a Together AI customer, noted they previously relied on manual versioning; the new platform lets them canary fine-tuned models on live traffic weekly, automatically rolling back if key metrics regress. Architects evaluating open-source inference should assess whether platform features like auto-rollback and safe iteration velocity outweigh the operational cost of self-hosting or the lock-in trade-off of closed-model APIs.

Sources