Hugging Face and alphaXiv orchestrated a 19-day effort with 1,221 community members armed with coding agents to reproduce papers from ICML 2026. The result is the largest attempted post-publication audit of a major ML conference: 6,816 Trackio logbooks, 2,226 papers attempted (34% of the conference), 35,908 claims judged, and a public searchable dataset. ICML 2026 accepted 6,352 papers—roughly double the prior year—from 23,918 submissions while reviewer capacity stalled.

MetricValue
Duration19 days
Community participants1,221
Papers attempted2,226 (34 % of conference)
Trackio logbooks produced6,816
Total claims judged35,908
ICML 2026 accepted papers6,352
ICML 2026 total submissions23,918
Top individual reproductions (one entrant)360+
Compute credit reward per participant$20
Top prize$2,000
FIG. 02 ICML 2026 Open Reproduction Challenge — top-level audit statistics — Hugging Face, huggingface.co/blog/icml-2026-open-reproductions

Participants used Claude Code, Codex, Cursor, or OpenResearch's orx to write and execute reproduction code on Hugging Face's serverless GPU compute. Each run produced a Trackio logbook—a Hugging Face Space with write-up, code, artifacts, and optional agent traces. An automated judge running GLM-5.2 assigned per-claim verdicts: verified, falsified, toy-scale, or inconclusive. Participants earned $20 in compute credits; the top prize was $2,000 out of $4,000. One entrant reproduced 360+ papers.

End-to-end reproduction pipeline: from participant to per-claim verdict
FIG. 03 End-to-end reproduction pipeline: from participant to per-claim verdict — Hugging Face, huggingface.co/blog/icml-2026-open-reproductions

51% of examined papers (1,103) had at least one claim independently verified. Of those, 266 were fully reproduced and 632 partially reproduced with no claims falsified—totaling 3,978 confirmed claims backed by real experiments. 23% of papers (496) had at least one claim falsified or contested. The category counts overlap because papers landing in both groups (verified claims and falsified claims in the same paper) count twice. The 242 papers receiving opposite verdicts from independent teams are the clearest case. 49 papers had all claims falsified and none verified.

Breakdown of the 2,226 attempted papers by reproduction outcome (categories may overlap)
FIG. 04 Breakdown of the 2,226 attempted papers by reproduction outcome (categories may overlap) — Hugging Face, huggingface.co/blog/icml-2026-open-reproductions

502 papers yielded only toy-scale evidence when datasets were proprietary or checkpoints unreleased. 280 papers produced nothing because artifacts were missing. Common failure modes: incomplete code, complex dependency chains, GPU requirements unavailable on commodity hardware, proprietary data with no synthetic stand-in. Where full reproduction was impossible, participants documented why and ran reduced-scale analogs, preserving the audit chain.

Two papers survived maximum scrutiny. "Flat Minima and Generalization: Insights from Stochastic Convex Optimization" was reproduced by 20 independent teams; 12 verified every claim. "A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness"—testing LLM judges themselves—had 14 of 17 logbooks verify all claims. At the other end, one accepted spotlight paper had a reviewer note stating they had not checked proofs carefully. The audit did. The falsification write-up is public.

Paper (abbreviated)Independent teamsTeams verifying ALL claims
Flat Minima and Generalization: Insights from Stochastic Convex Optimization2012
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness1714
FIG. 05 Papers with the most independent reproduction teams and their all-claims verification rates — Hugging Face, huggingface.co/blog/icml-2026-open-reproductions

Trackio installs via pip. The scaffold script generates a logbook skeleton with a single curl command. The 35,908-claim verdict dataset is public on Hugging Face. The challenge protocol—extract claims, run experiments against each claim, have an independent judge re-read logbooks with self-assessments marked untrusted—is a pattern platform teams can adopt for internal model evaluation. The 242 papers with contradictory verdicts form the primary product: a ranked, queryable list of contested claims in current ML research, with agent traces attached.