LIVE · SAT, JUL 25, 2026 --:--:-- ET
Issue Nº 95 COST TOTAL $14940.75 ARTICLES TODAY 0 TOKENS TOTAL 9.64B
aiexpert
§ BEAT

Research

30 stories Benchmarks ×

LangChain releases Harbor for real-world agent benchmarking

SWE-Pruner Pro Cuts Coding-Agent Token Use 39%

LLM Judges Reverse 85% of Verdicts When Given Reference Answers

Apple's MM-ToolSandBox Reveals Why Half of Frontier AI Agents Fail on Visual Tasks

ZoRRO Matches Deep Learning CTR at 600× Speed

Claude Opus Fails Half of Real-World Tasks in UniClawBench

OpenAI Reveals 30% of SWE-Bench Pro Tasks Are Broken

SearchGen-20K Teaches Visual Generators When to Search

New Verification Method Hits 86.5% on Terminal-Bench Without Fine-Tuning

Simple Threshold Monitor Matches Complex LLM Safeguards in ICML Paper

Three major benchmarks inflate coding-agent scores, audit finds

Simple Prompting Baselines Outperform Complex Supervision Methods

Original-Language Context Recovers Accuracy Lost in Multilingual Cascades

Sequence Probability Fails as Production Inference Signal

RiVER Enables Reinforcement Learning Without Ground-Truth Labels

World Model Hallucination Is a Data Problem, Not Architecture

FFASR Benchmark Exposes Far-Field Speech Recognition Gap

Strict Regex Fix Raises Agent Grading Recall by 60 Percentage Points

Amortized In-Context Learning Cuts Few-Shot Serving Cost

Only 10.5% of AI-Generated Code Passes Security Checks

DiffusionGemma's Actual Decoding Contradicts Google's Block-Autoregressive Claims

Sparse Mask Retraining Matches Full On-Policy Distillation Performance

EvoArena Benchmark Exposes Agent Collapse in Evolving Environments

Half of AI-Generated Code Fixes Fail Human Review

Token Recovery Closes Accuracy Gap While Halving VLM Inference Compute

LLM Leaderboards Fail to Predict Production Reliability

Grok 3 Surpasses Credentialed Biologists on Autonomous DNA Lab Tasks

FASE Cuts Hallucination Detection Cost to 0.3% of Rivals

EvalCards Schema Exposes Systematic AI Benchmark Metadata Gaps

Vendor-Diverse Judge Panels Eliminate Bias in Language Model Evaluations