Everything the newsroom published, in chronological order. Each item carries origin, sources and reading time.
RESEARCH Self-Modifying Agents Boost Benchmark Score to 0.61
RESEARCH DeltaBox cuts AI agent checkpoint latency to 14 milliseconds
RESEARCH LCGuard Patches KV-Cache Leakage in Multi-Agent Systems
RESEARCH DelTA Framework Improves Reasoning by Fixing Token-Level Credit Assignment
RESEARCH Equilibrium Reasoners lift Sudoku accuracy from 2.6% to 99% via test-time scaling
RESEARCH NVIDIA's CARV cuts 3D distillation compute by 2–3×
RESEARCH One hyperparameter rule captures most of µP's gains
RESEARCH RELEX reconstructs RLVR checkpoints from 15% training data
RESEARCH Peking researchers release DeepWeb-Bench, exposing derivation failures in frontier AI
RESEARCH Fine-tuning erases reasoning chains while accuracy stays high
RESEARCH Allen AI's OlmoEarth v1.1 cuts satellite inference compute 3x
RESEARCH Medical LLMs Underweight Patient Autonomy
RESEARCH Researchers Map Hallucination Rates by Model Size and Data Frequency
RESEARCH RRFP Achieves 2.77× Throughput on Multimodal Pipeline-Parallel Training
RESEARCH DashAttention reaches 75% sparsity while matching full-attention accuracy
RESEARCH EnvFactory lifts Qwen3 tool-calling accuracy 15% with synthetic data
RESEARCH Memory Lookup Replaces Linear Attention Over Long Prefixes
RESEARCH SAEBench Metrics Rank SAEs Backwards, Audit Finds
RESEARCH Autonomous Disease Forecasting System Outperforms CDC Ensemble on Blinded Tests
RESEARCH FORGE Reduces Agent Failures to 1% Without Model Fine-Tuning
RESEARCH Grep Beats Vector Search in Inline Agent Retrieval
RESEARCH Frontier Agents Reach 25% on Real-World Forecasting Test
RESEARCH Microsoft Finds GPT-5 Fails Against Implausible Attacks
RESEARCH Scientific ML Models Disagree on 16% of Predictions Despite Matching Accuracy
RESEARCH LLM Formalization Catches 18.8% Ambiguous Requirements in Safety Specs
RESEARCH TFlow cuts multi-agent inference tokens 83% via weight injection
RESEARCH Negation Neglect Drives False Belief Rate to 88.6% in Fine-Tuned LLMs
RESEARCH Why Production Agents Fail Without Harness Infrastructure
RESEARCH Berkeley Framework Cuts Agent Latency 1.3–2.2×
RESEARCH KV-Fold Extends Transformer Context to 128K Without Retraining
RESEARCH IBM Boosts Zero-Shot Search Accuracy 25% With LLM Query Refinement
RESEARCH 27M Attractor Model Beats GPT o3 on Logic Puzzles
RESEARCH Reward Hacking Undetected in Single-Verifier Training
RESEARCH Sparse-to-Dense RL Lifts MATH Scores to 78.5% on Small Models
RESEARCH Standard load-balancing losses degrade SMoE expert specialization by 3x