Everything the newsroom published, in chronological order. Each item carries origin, sources and reading time.
RESEARCH Super Weights Training Fails on OLMo Models, Demolishing Sparse Fine-Tuning Strategy
RESEARCH Training-Efficient Low-Rank Compression Sidesteps Serving-Speed Proof
RESEARCH Hugging Face Cuts Inference Attention Overhead 20-40% With Fused Kernels
RESEARCH Claude Opus Fails Half of Real-World Tasks in UniClawBench
RESEARCH Cornell's Co-LMLM Matches GPT-4o-Mini by Storing Facts in a Database RESEARCH Timestep Weighting Cuts Reward-Model Query Costs for Diffusion RLHF
RESEARCH STRACE Framework Boosts Multi-Agent Verification by 16 Points
RESEARCH DynaKRAG Boosts Multi-Hop QA Accuracy by Up to 5.78 Points
RESEARCH OpenAI Reveals 30% of SWE-Bench Pro Tasks Are Broken
RESEARCH DepthWeave-KV cuts LLM cache memory by 8.3x without retraining
RESEARCH SovereignPA-Bench Measures Whether AI Agents Protect User Boundaries
RESEARCH Dual Planning Loop Solves Semantic Gap in Hierarchical Robots
RESEARCH SearchGen-20K Teaches Visual Generators When to Search
RESEARCH CompactionRL boosts GLM coding agents 5–7 points on benchmarks
RESEARCH New Verification Method Hits 86.5% on Terminal-Bench Without Fine-Tuning
RESEARCH Simple Threshold Monitor Matches Complex LLM Safeguards in ICML Paper
RESEARCH Misaligned Coding Agents Evade Monitors in 93% of Gradual Attacks
RESEARCH LACUNA Shows Unlearning Methods Fail to Erase PII from Models
RESEARCH Language Labels Beat Scalars in Offline Robot Learning
RESEARCH Theoria bridges formal proof and LLM judges with auditable verification
RESEARCH Three major benchmarks inflate coding-agent scores, audit finds
RESEARCH AutoMem Training Doubles Agent Performance on Long-Horizon Tasks
RESEARCH One Layer Matches Full RL Post-Training on Qwen Models
RESEARCH BrowserBC Lifts Browser Agent Success to 81% Using Human Traces
RESEARCH Language Model Explanations Track Behavior Shifts Automatically
RESEARCH TRIAGE Cuts Agent Actions 14.8% While Raising Success Rates
RESEARCH Simple Prompting Baselines Outperform Complex Supervision Methods
RESEARCH Researchers Close Gap Between AI Agents and Hand-Curated Skills
RESEARCH New Training Technique Improves LLM Confidence Calibration by 63%
RESEARCH Google Releases Zero-Shot Tabular Model but Hides Benchmark Data
RESEARCH Vision-language models route knowledge through just 2.5% of network RESEARCH AI Agents Double Repository-Level Merge Friction
RESEARCH Original-Language Context Recovers Accuracy Lost in Multilingual Cascades RESEARCH ENS Hits 10× Accuracy on Tough PDE Benchmarks Without Correction Loops
RESEARCH Mechanism Taxonomy Lifts LLM Moderation F1 by 5.4%