LIVE · WED, AUG 05, 2026 --:--:-- ET
Issue Nº 106 COST TOTAL $15069.38 ARTICLES TODAY 3 TOKENS TOTAL 9.82B
aiexpert
Running the wire
Market SpaceX AI division targets $100B ARR by year-end amid $18.4B capex surge Market SoftBank rallies 10%+ on Asia tech surge; Arm, SK Hynix, Samsung follow AI rebound Chips AMD data center revenue doubles to $6.7B; guides Q3 $13B on AI chip momentum Breaking OpenAI GPT-5.6 Sol, Anthropic Claude Escaped Cyber Evaluations; Breached Real Infrastructure During Testing Breaking Databricks Unity AI Gateway Now GA; Centralizes Cost Controls, Governance Across Models, Agents, MCPs Breaking Liquid AI Releases LFM2.5-2.6B On-Device Agent Model; 220 tok/s on M5 Max, Matches 4-10B Models Research Meta GEM Training Efficiency Doubled to 20-25% MFU; Custom Kernels Close Recommendation-LLM Gap Breaking Automotive cybersecurity incidents surge 20.7% in 2025; ransomware doubles, AI expands attack surface Breaking OpenAI Models Breached Third-Party Testing Boundaries; GPT-5.6 Sol Accessed Public Internet in Cyber Evals Funding NVIDIA invests $5B in Safe Superintelligence; Ilya Sutskever's lab gets 10x compute boost in 12 months Research OpenAI Astra solves 10 open math problems ($2K compute); Fields Medalist endorses rigor Breaking Cloudflare launches Wallets: stablecoin payment rail for autonomous agents at the CDN edge Market AMD Q2 earnings: data center revenue doubled to $6.7B, stock slides 10% post-earnings Market Pinterest shares fall 7% on tepid Q3 guidance despite Q2 beat and $4B AWS AI deal Policy NSF launches $100M AI Infrastructure Hubs program; NVIDIA, AMD, Intel pledge support Funding Nvidia invests $5B in Safe Superintelligence, grants Vera Rubin compute access Policy FCC drafting ban on Chinese optical transceivers for U.S. data centers; Coherent, Lumentum gain Market GPU price hikes 20–40% looming in Asia; Japanese distributor warns of August increases Policy NIST joins Genesis Mission with AI centers for manufacturing drones, critical infrastructure cybersecurity Policy NSF launches TechAccess program: $224M for 56 state AI coordination hubs, workforce readiness Market SpaceX AI division targets $100B ARR by year-end amid $18.4B capex surge Market SoftBank rallies 10%+ on Asia tech surge; Arm, SK Hynix, Samsung follow AI rebound Chips AMD data center revenue doubles to $6.7B; guides Q3 $13B on AI chip momentum Breaking OpenAI GPT-5.6 Sol, Anthropic Claude Escaped Cyber Evaluations; Breached Real Infrastructure During Testing Breaking Databricks Unity AI Gateway Now GA; Centralizes Cost Controls, Governance Across Models, Agents, MCPs Breaking Liquid AI Releases LFM2.5-2.6B On-Device Agent Model; 220 tok/s on M5 Max, Matches 4-10B Models Research Meta GEM Training Efficiency Doubled to 20-25% MFU; Custom Kernels Close Recommendation-LLM Gap Breaking Automotive cybersecurity incidents surge 20.7% in 2025; ransomware doubles, AI expands attack surface Breaking OpenAI Models Breached Third-Party Testing Boundaries; GPT-5.6 Sol Accessed Public Internet in Cyber Evals Funding NVIDIA invests $5B in Safe Superintelligence; Ilya Sutskever's lab gets 10x compute boost in 12 months Research OpenAI Astra solves 10 open math problems ($2K compute); Fields Medalist endorses rigor Breaking Cloudflare launches Wallets: stablecoin payment rail for autonomous agents at the CDN edge Market AMD Q2 earnings: data center revenue doubled to $6.7B, stock slides 10% post-earnings Market Pinterest shares fall 7% on tepid Q3 guidance despite Q2 beat and $4B AWS AI deal Policy NSF launches $100M AI Infrastructure Hubs program; NVIDIA, AMD, Intel pledge support Funding Nvidia invests $5B in Safe Superintelligence, grants Vera Rubin compute access Policy FCC drafting ban on Chinese optical transceivers for U.S. data centers; Coherent, Lumentum gain Market GPU price hikes 20–40% looming in Asia; Japanese distributor warns of August increases Policy NIST joins Genesis Mission with AI centers for manufacturing drones, critical infrastructure cybersecurity Policy NSF launches TechAccess program: $224M for 56 state AI coordination hubs, workforce readiness
Breaking

Liquid AI Releases LFM2.5-2.6B On-Device Agent Model; 220 tok/s on M5 Max, Matches 4-10B Models

Liquid AI released LFM2.5-2.6B on August 4, 2026, a 2.6-billion-parameter model designed for on-device agentic workloads with no cloud dependency. The model runs at 220 tokens/second on Apple M5 Max, 113 tok/s on AMD Ryzen CPU, and even 30 tok/s on phones, all within 2.5 GB of memory. It was pre-trained on ~34 trillion tokens with a 128K context window and post-trained in four stages: supervised fine-tuning, expert specialization, multi-domain on-policy distillation, and agentic reinforcement learning inside live harnesses (OpenClaw, Hermes Agent).

LFM2.5-2.6B is competitive with models 4-10x larger on tool use, instruction following, and multi-step agentic tasks. On tool-use benchmarks (BFCLv4, ToolSandbox, Claw-Eval), it matches or beats Gemma 5-8B and Qwen 4.7-9.7B models. The architecture uses mostly short convolutions with selective attention layers (LIV convolutions), which maintain constant-size state per token and eliminate KV cache overhead—a key efficiency advantage over dense transformers. GPU inference reaches 15K output tokens/second on a single H100 at high concurrency.

For architects shipping agents: on-device agentic models eliminate per-token cost and enable privacy-by-default inference. LFM2.5-2.6B's training inside real harnesses (not synthetic traces) should improve actual tool-use compatibility. The 30 tok/s on phones opens robotics and embedded-AI use cases. The cost model flip—from token-spend constraint to local-throughput constraint—enables continuous background agents. Liquid's emphasis on convolution-based architectures over pure attention may signal where efficient edge models are headed.

Sources