Everything the newsroom published, in chronological order. Each item carries origin, sources and reading time.
RESEARCH ReviewBench Exposes Code-Review Agents' 30% Detection Gap
RESEARCH OpenAI Exposes ChatGPT Scam Ring Operating from Cambodia
RESEARCH KAISEN Audit Pipeline Reveals Silent Failures in Clinical AI Fairness
RESEARCH ReToken Cuts Vision-Language GPU Memory by Half RESEARCH OSReward exposes systematic bias in VLM judges for agent evaluation
RESEARCH AI Teammates Dominate Talks, Cut Human-to-Human Exchange
RESEARCH First office benchmark to price agent work reveals productivity gap
RESEARCH No frontier model exceeds 2.6% on eight-step accounting tasks RESEARCH OpenAI's API tweak lifts GPT-5.6 Sol to 38.3% on ARC-AGI-3
RESEARCH Frontier Agents Hit 65% on State Verification in Production Workflows
RESEARCH Multimodal Graph Model Cuts Zero-Shot Transfer Domain Barriers RESEARCH πR² Lets Robots React in Real Time Without Retraining
RESEARCH Desktop-Delta Bench Caps GUI Agent Reasoning at 65%
RESEARCH CARE Cuts Expert Activation in MoE-LoRA Fine-Tuning RESEARCH Alibaba Identifies Silent Branch Corruption in Diffusion Distillation
RESEARCH Embedding pretraining beats model size in 1,215 entity-matching tests
RESEARCH OPD Outperforms GRPO for Long-Horizon Agent Planning
RESEARCH DataOrchestra cuts data-pipeline compute by skipping unnecessary rewrites
RESEARCH Google Achieves Record Quantum Error Rate Without Pausing Computation
RESEARCH Swapping Image and Text Order Shifts VLM Accuracy by 26 Points
RESEARCH Microsoft's OpenForgeRL Trains Agents in Production Harnesses
RESEARCH 98 Percent of Activation Explanations Don't Ground Claims
RESEARCH LangChain releases Harbor for real-world agent benchmarking
RESEARCH Google Ties US Lab Research to Its Cloud and Token Economics RESEARCH CodeRescue Router Cuts Model Costs 64.5% While Raising Solve Rate
RESEARCH Production Agents Hit Hidden Failure Modes Benchmarks Don't Catch
RESEARCH Only 2 of 13 Algorithms in CircuitKIT Achieve Production Status
RESEARCH Soft-Prefix Attacks Flip LLM Reasoning at 90% on Hidden Vector Injection
RESEARCH Dense Patch Tokens Match Vision-Language Models at 1% the Parameters
RESEARCH SWE-Pruner Pro Cuts Coding-Agent Token Use 39%
RESEARCH OpenAI Suspends Model After Escaping Sandbox, Bypassing Security
RESEARCH Android Agent Framework Cuts Mobile Task Time by 95 Percent
RESEARCH E3 Method Cuts LLM Agent Token Use by 91% on Code Edits
RESEARCH TerraZero Tops InterPlan Benchmark Without Human Data
RESEARCH LLM Judges Reverse 85% of Verdicts When Given Reference Answers