News
AI, at newsroom pace.
Anthropic ships Claude Opus 5 at $5/$25 per Mtok—near-Fable-5 coding at half Fable's price
RESEARCH
Anthropic releases Claude Opus 5: matches Fable on benchmarks, 50% cheaper, praised for coding
RESEARCH
Kimi K3 achieves 68.5% on DeepSWE, costs 2.8x less per task than Claude Fable 5
RESEARCH
Black Forest Labs FLUX 3: unified multimodal model generates video, audio, and robot actions from single architecture
RESEARCH
Microsoft's OpenForgeRL Trains Agents in Production Harnesses
RESEARCH
NVIDIA research at ICML 2026: 145 papers cite Nemotron open models; 2,000 papers use NVIDIA GPUs
RESEARCH
98 Percent of Activation Explanations Don't Ground Claims
RESEARCH
LangChain releases Harbor for real-world agent benchmarking
RESEARCH
Poolside releases Laguna S 2.1: 118B open-weight coding model matches models 10x larger, trains in under 9 weeks
RESEARCH
Production Agents Hit Hidden Failure Modes Benchmarks Don't Catch
RESEARCH
Poolside releases Laguna S 2.1, 118B open-weight model beating DeepSeek V4 Flash on agentic tasks
RESEARCH
LangChain publishes Deep Agents benchmark framework; three eval suites (Harbor-Index, τ³-bench, ContextBench) set long-horizon agent autonomy standard
RESEARCH
Nvidia launches DNA genomics model; learns what token prediction misses in biological data
RESEARCH
Moonshot Kimi K3: Chinese open-weight model tops Arena benchmark, outranks Claude on code
RESEARCH
Poolside releases Laguna S 2.1, 118B-parameter open-weight model matching closed competitors
RESEARCH
Nunchaku brings SVDQuant 4-bit diffusion inference to Hugging Face, Diffusers ecosystem
RESEARCH
Google Ties US Lab Research to Its Cloud and Token Economics
RESEARCH
CodeRescue Router Cuts Model Costs 64.5% While Raising Solve Rate
RESEARCH
Only 2 of 13 Algorithms in CircuitKIT Achieve Production Status