§ BEAT
Research
98 Percent of Activation Explanations Don't Ground Claims
Only 2 of 13 Algorithms in CircuitKIT Achieve Production Status
Activation-Level Fixes Outperform Prompt Edits for Biased LLM Judges
Super Weights Training Fails on OLMo Models, Demolishing Sparse Fine-Tuning Strategy
LACUNA Shows Unlearning Methods Fail to Erase PII from Models
Language Model Explanations Track Behavior Shifts Automatically
Vision-language models route knowledge through just 2.5% of network
Models Shed Learned Rules During Training
Multimodal Models Flip Answers When Evidence Order Changes
Google DeepMind's DiffusionGemma 28.6X harder to interpret than autoregressive models
MIT Extracts Attention Logic Into Swappable Python Code
Sparse Attention Heads Redirect Vision-Language Models With 83% Accuracy
Label-Free Test Catches LLM Reasoning Failures Better Than Self-Consistency
New Tool Finds 1,060 Hidden Training Dependencies Across Major LLMs
Kamai's Phase Diagram Predicts Multimodal Failure Before GPU Commit
Real EHR Benchmark Exposes Limits of LLMs in Clinical Action
Echo-Memory Shows World Models Fail the Revisit Test
64 Percent of Audio-Text Conflicts in AI Models Are Fixable
Stanford Framework Keeps AI Agents Within Violation Targets
Self-Generated Replay Cuts Catastrophic Forgetting in Fine-Tuned Models
Study: AI Narrative Explanations Boost User Trust, Not Accuracy
DelTA Framework Improves Reasoning by Fixing Token-Level Credit Assignment
RELEX reconstructs RLVR checkpoints from 15% training data
SAEBench Metrics Rank SAEs Backwards, Audit Finds
Math Proof Shows Transformer Attention Stabilizes Predictably
SLIM improves LLM agent performance 7 percentage points
Shepherd Raises Agent Accuracy 90% With Forking Traces
Sparse MoE Models Match Dense Transformers at 3× Faster Inference
Frozen Models Encode Semantic Roles Without Fine-Tuning