Wire
Accounting agents top out at 56.4% on APEX
Mercor and Ramp’s APEX-Accounting benchmark gives frontier agents 160 professional tasks across 10 simulated work environments, but the leader reached only 56.4% on its partial-credit metric. The benchmark preprint reports that no tested model exceeded 2.6% on eight-run consistency and the best model completed every rubric item in at least one of eight attempts on just 21.5% of tasks, despite token budgets rising from $1 to $50. For finance teams evaluating enterprise agent platforms as governed workflow infrastructure, the gap between average credit and repeatable completion argues for rubric-level evidence and human approval before any agent posts transactions or changes a ledger.