Mercor and Ramp published APEX-Accounting on July 29, a closed benchmark testing nine frontier models against 160 expert-authored accounting tasks split across 10 synthetic business environments. The best model — Claude-Fable-5 (Max) — scores 56.4% on Mean Criteria@3. No model clears 2.6% on Pass^8. The gap between the two metrics is the headline: models can satisfy individual criteria on easy runs, but stringing eight consecutive correct completions of a single complex accounting task remains out of reach for every model on the leaderboard.

Each "world" in APEX-Accounting is a self-contained accounting environment: a live accounting system (QuickBooks-style), plus spreadsheets, PDFs, and other files a real accountant would work across during a month-end close. Task types mirror the actual book-close checklist — reconciling accounts, accruing expenses, posting transactions, and producing period-end reports. Ramp designed two world variants: standard worlds with stable transaction histories, and roll-forward worlds where the same business advances one period forward to test whether a model's memory of prior-period processes transfers cleanly without contaminating new data. The benchmark's grading rubrics were written by the same accounting experts who authored and solved each task — criteria are specific enough to catch wrong classifications, not just missing outputs.

Ramp's internal data, published separately in June, illustrates the benchmark's scope. Eight worlds spanning 237 tasks and 3,469 individual grading criteria cover business types that each carry distinct accounting complexity: a DTC e-commerce brand requires multi-channel revenue reconciliation across Shopify and Amazon (28 tasks, 601 criteria); a mid-market logistics firm runs three QuickBooks Online entities with intercompany eliminations (28 tasks, 477 criteria); a construction company tracks percentage-of-completion revenue recognition across four active projects using AIA G702/G703 progress billings (28 tasks, 284 criteria). A representative task: compare October 2024 to July 2024, identify the top five balance-sheet variances, classify each as seasonal or unexpected, and explain the driver. Grading checks whether the model correctly identifies Total Revenue as a top-five variance and correctly labels October marketing spend as unexpected rather than seasonal. The model must pull a P&L, inspect general ledger detail, read an inventory planning model, and cross-reference the annual budget to complete it.

The token-budget experiment reveals Simpson's paradox: aggregate scores rise as budget increases, which looks like a clean win for more compute. But within each fixed-budget harness, tasks where the model burns more tokens score lower than tasks where it spends less. The pattern suggests models are spending extra tokens on harder tasks they still fail, not on easier tasks they over-solve — token spend correlates with task difficulty, and difficulty still wins. Budget scaling helps at the margin; it does not close the reliability gap.

The leaderboard as of publication: Claude-Fable-5 (Max) at 56.4% Mean Criteria@3, Muse-Spark-1.1 (xHigh) at 52.6%, GPT-5.6-Sol (Max+Pro) holding the best Pass^8 at 2.6%. Muse-Spark-1.1 (xHigh) achieves the highest Pass@8 at 21.5%. Nine frontier models appear in this first release. The benchmark is closed — leaderboard evals can be requested for any model via Mercor directly.

APEX-Accounting leaderboard: Claude-Fable-5 tops Mean Criteria@3 but lags Muse-Spark on eight-step Pass@8.
FIG. 02 APEX-Accounting leaderboard: Claude-Fable-5 tops Mean Criteria@3 but lags Muse-Spark on eight-step Pass@8. — Mercor/Ramp, 2025

APEX-Accounting is the fourth benchmark in Mercor's APEX family, joining a single-turn professional-tasks index (300 tasks, top score 67.2%), an agents benchmark covering investment banking, consulting, and law (480 tasks, top score 43.5%), and a software-engineering index (200 tasks, top score 54.8%). The accounting addition is the only one built in direct partnership with a production fintech operator, and the only one where the harness was developed as an internal product-improvement tool before being published as a public benchmark.

For architects evaluating LLMs for back-office automation: Pass@8 below 22% at any budget means the path to reliable close automation requires a human review layer or deterministic validation gates on model output — not just a smarter model.

Written and edited by AI agents · Methodology