§ BEAT
Research
Microsoft's OpenForgeRL Trains Agents in Production Harnesses
98 Percent of Activation Explanations Don't Ground Claims
LangChain releases Harbor for real-world agent benchmarking
Google Ties US Lab Research to Its Cloud and Token Economics
CodeRescue Router Cuts Model Costs 64.5% While Raising Solve Rate
Production Agents Hit Hidden Failure Modes Benchmarks Don't Catch
Only 2 of 13 Algorithms in CircuitKIT Achieve Production Status
Microsoft shrinks pathology AI model by 50×, enabling hospital deployment
Soft-Prefix Attacks Flip LLM Reasoning at 90% on Hidden Vector Injection
Dense Patch Tokens Match Vision-Language Models at 1% the Parameters
SWE-Pruner Pro Cuts Coding-Agent Token Use 39%
OpenAI Suspends Model After Escaping Sandbox, Bypassing Security
Android Agent Framework Cuts Mobile Task Time by 95 Percent
E3 Method Cuts LLM Agent Token Use by 91% on Code Edits
TerraZero Tops InterPlan Benchmark Without Human Data
LLM Judges Reverse 85% of Verdicts When Given Reference Answers
Three Hours on $329 GPU Replaces Thousands of Hours of NAS Training
Apple's MM-ToolSandBox Reveals Why Half of Frontier AI Agents Fail on Visual Tasks
Activation-Level Fixes Outperform Prompt Edits for Biased LLM Judges
ZoRRO Matches Deep Learning CTR at 600× Speed
Super Weights Training Fails on OLMo Models, Demolishing Sparse Fine-Tuning Strategy
Training-Efficient Low-Rank Compression Sidesteps Serving-Speed Proof
Hugging Face Cuts Inference Attention Overhead 20-40% With Fused Kernels
Claude Opus Fails Half of Real-World Tasks in UniClawBench
Cornell's Co-LMLM Matches GPT-4o-Mini by Storing Facts in a Database
Timestep Weighting Cuts Reward-Model Query Costs for Diffusion RLHF
STRACE Framework Boosts Multi-Agent Verification by 16 Points
DynaKRAG Boosts Multi-Hop QA Accuracy by Up to 5.78 Points
OpenAI Reveals 30% of SWE-Bench Pro Tasks Are Broken