APEX-Accounting: An AI Productivity Benchmark for Accounting Work
Key point
APEX-Accounting evaluated whether AI can repeatedly produce accurate results in accounting tasks.
Details
APEX-Accounting is an accounting AI agent benchmark jointly developed by Ramp and Mercor. It evaluates 160 tasks written by accounting experts across 10 simulated company environments, measuring reconciliation of discrepancies between files, application of company-specific context, multi-step judgment, and consistency of results.
Each model performed every task 8 times. Because the benchmark values repeated success over a single correct answer, even the most consistent model succeeded on all 8 runs for only 2.6% of all tasks. While tasks where a model earned at least partial credit exceeded 95%, tasks that no model ever fully solved reached as high as 58%.
On the average pass rate, Claude Fable 5 ranked first at 56.4%, followed by Meta's Muse Spark 1.1 at 52.6% and GPT-5.6 Sol at 51.5%. On Pass@8 (succeeding at least once in 8 tries), Muse Spark 1.1 led with 21.5%, ahead of Fable 5's 20.1%.
Cost budgets were set at $1, $5, $10, and $50 per task. Fable 5 scored 11.8% at the $1 budget but improved significantly to 55.2% at $50, whereas Muse Spark 1.1 already performed well even at lower budgets, showing only small gains from additional spending. At the $50 cap, Fable 5 used about $32 per run while Muse Spark 1.1 used only about $5, yet the score gap between the two models stayed within 4 percentage points.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.