Everything the newsroom published, in chronological order. Each item carries origin, sources and reading time.
RESEARCH Microsoft Maps Multilingual Agent Planning Failures, Lifts Accuracy 5.6%
RESEARCH GDPevo benchmark shows enterprise agents gain 8 points from experience
RESEARCH HazMart Quantifies Faithfulness-Safety Trade-Off in Reasoning Models
RESEARCH Frontier Models Attempt Cheating in 7–14% of Evaluation Runs
RESEARCH AtumAI Cuts Datacenter Policy Design from Months to Plain Text
RESEARCH PRECOG Cuts RAG Latency 4500× on Edge SSM Hardware
RESEARCH Caltech Robot Dodges Balls With 95% Accuracy Using Only a Camera
RESEARCH Smevals cuts model eval cost by decoupling grading from test runs RESEARCH Maryland team tunes OPSD's hidden constant to fix brittleness
RESEARCH Repeated Sampling Beats Self-Reflection on Small AI Models
RESEARCH ReviewBench Exposes Code-Review Agents' 30% Detection Gap
RESEARCH OpenAI Exposes ChatGPT Scam Ring Operating from Cambodia
RESEARCH KAISEN Audit Pipeline Reveals Silent Failures in Clinical AI Fairness
RESEARCH ReToken Cuts Vision-Language GPU Memory by Half RESEARCH OSReward exposes systematic bias in VLM judges for agent evaluation
RESEARCH AI Teammates Dominate Talks, Cut Human-to-Human Exchange
RESEARCH First office benchmark to price agent work reveals productivity gap
RESEARCH No frontier model exceeds 2.6% on eight-step accounting tasks RESEARCH OpenAI's API tweak lifts GPT-5.6 Sol to 38.3% on ARC-AGI-3
RESEARCH Frontier Agents Hit 65% on State Verification in Production Workflows
RESEARCH Multimodal Graph Model Cuts Zero-Shot Transfer Domain Barriers RESEARCH πR² Lets Robots React in Real Time Without Retraining
RESEARCH Desktop-Delta Bench Caps GUI Agent Reasoning at 65%
RESEARCH CARE Cuts Expert Activation in MoE-LoRA Fine-Tuning RESEARCH Alibaba Identifies Silent Branch Corruption in Diffusion Distillation
RESEARCH Embedding pretraining beats model size in 1,215 entity-matching tests
RESEARCH OPD Outperforms GRPO for Long-Horizon Agent Planning
RESEARCH DataOrchestra cuts data-pipeline compute by skipping unnecessary rewrites
RESEARCH Google Achieves Record Quantum Error Rate Without Pausing Computation
RESEARCH Swapping Image and Text Order Shifts VLM Accuracy by 26 Points
RESEARCH Microsoft's OpenForgeRL Trains Agents in Production Harnesses
RESEARCH 98 Percent of Activation Explanations Don't Ground Claims
RESEARCH LangChain releases Harbor for real-world agent benchmarking
RESEARCH Google Ties US Lab Research to Its Cloud and Token Economics RESEARCH CodeRescue Router Cuts Model Costs 64.5% While Raising Solve Rate
RESEARCH Production Agents Hit Hidden Failure Modes Benchmarks Don't Catch
RESEARCH Only 2 of 13 Algorithms in CircuitKIT Achieve Production Status