Everything the newsroom published, in chronological order. Each item carries origin, sources and reading time.
RESEARCH Unified LLM Training Reveals Fundamental Mode Conflict
RESEARCH MMDiff Lets Engineers Audit and Edit Multimodal Model Features
RESEARCH ArchAgent v2 beats human champions on three-level cache design RESEARCH GENCO replaces three classical power grid solvers with one neural architecture
RESEARCH Amazon's Consilience Detects Silent Mode Collapse in Confidence-Based Reasoning
RESEARCH Taboo Stress-Tests LLMs on Production Constraints
RESEARCH Muon Optimizer Collapse After Grokking Threatens Production Training
RESEARCH SABRE Benchmark Exposes Vision Models' Blindness to Visual Contradictions
RESEARCH CreativeInstruct Recovers Diversity Without Slowing Inference
RESEARCH Mankind Labs Cuts Agentic Loop Tokens 17–26% With Reversible Memory Eviction RESEARCH CoinRAG Achieves 5.3% F1 Gain on Multi-Hop RAG with Nugget Caching
RESEARCH CalibForge Trains Models to 30-Point Gains with Solver-Calibrated Tasks
RESEARCH AV-AIVAT Cuts Agent Evaluation Cost 74-Fold With Certified Stopping RESEARCH Programmatic Tool Calling Beats JSON on New AI Models RESEARCH NVIDIA Publishes Full Greek RAG Playbook Reaching 2.3× Retrieval Gain
RESEARCH DeepMind's WeatherNext Cyclones Outpace Hurricane Models by Full Day
RESEARCH Skill Entropy Lifts Reasoning Model Scores From 34% to 68%
RESEARCH OctoLong improves code agents with cross-repository training
RESEARCH EvolveNet Evolves Agent Harness Instead of Model Weights
RESEARCH Chain-of-Thought Works at Tiny Scale, Breaking Compute Link RESEARCH SeGaBench Tests LLMs on Compiler-Blind Code Optimization
RESEARCH ReflectRL Recovers Reasoning Gains from Failed Expert Rollouts
RESEARCH Test-time scaling regimes require distinct evaluation profiles RESEARCH ALiBi Models Lose Positional Awareness in Long-Context Inference
RESEARCH Launch Inputs Matter More Than Model Choice for TUI Testing
RESEARCH Microsoft Maps Multilingual Agent Planning Failures, Lifts Accuracy 5.6%
RESEARCH GDPevo benchmark shows enterprise agents gain 8 points from experience
RESEARCH HazMart Quantifies Faithfulness-Safety Trade-Off in Reasoning Models
RESEARCH Frontier Models Attempt Cheating in 7–14% of Evaluation Runs
RESEARCH AtumAI Cuts Datacenter Policy Design from Months to Plain Text
RESEARCH PRECOG Cuts RAG Latency 4500× on Edge SSM Hardware