Ramp and Mercor built APEX-Accounting, a benchmark that runs AI models through 160 tasks spread across 10 simulated companies. Closing a company’s books is not a single-answer problem: an agent must reconcile conflicting source files, hold onto context specific to that business, and keep a conclusion consistent as it moves through several steps of the workflow. Mercor argues that a model which drops a correct intermediate finding can still generate a bad journal entry, which is why it tested every model eight separate times per task rather than once.

That repetition produced the benchmark’s starkest number. The most consistent model, Claude Fable 5, solved only 2.6 percent of tasks correctly across all eight runs. On the standard leaderboard, which averages performance rather than demanding perfect consistency, Fable 5 topped the field at 56.4 percent. Muse Spark 1.1, Meta’s entry, followed at 52.6 percent, with GPT-5.6 Sol close behind at 51.5 percent. Qwen 3.5 finished lowest at 24.4 percent. Partial credit was common (more than 95 percent of tasks earned at least some credit from some model), but full, correct solutions were rare, and 58 percent of the tasks went completely unsolved by every model across every run.

Spending more money did not reliably close that gap. Mercor capped each model’s spending at four separate levels, $1, then $5, then $10, then a top tier of $50, and reran the tasks at each. Fable 5 was highly budget sensitive, rising from 11.8 percent at the $1 cap to 55.2 percent at $50. Muse Spark 1.1 moved the other direction, already competitive when spending was capped and only marginally better as the budget grew. At the $50 ceiling, Fable 5 spent roughly $32 per run against Muse Spark’s roughly $5, yet the two finished within four percentage points of each other.

Mercor’s error analysis, built with accounting and bookkeeping experts, traced the gap to judgment rather than search. About seven in ten failures, across the top three models, traced back to flawed reasoning rather than missing information: a model would flag a discrepancy correctly early in a task, then contradict or drop that finding by the time it wrote the final journal entry.

The benchmark’s credibility rests on who wrote and graded its 160 tasks. Mercor says over 40 practicing accounting professionals wrote and completed the tasks themselves before anyone graded them, a group with a track record measured at a median of eleven years, more than half of them alumni of a Big Four firm. The specific tasks and their grading rubrics were built by experts with prior experience at KPMG, EY, PwC, and Deloitte, hired through Mercor’s marketplace, and each task carried a rubric averaging 13.7 criteria written by those same experts. Grading the runs, though, was not manual: an AI judge, released open source, evaluated the outputs and matched expert human graders 97 percent of the time, a figure Mercor does not break down by task type. The benchmark also covers only month-end close and bookkeeping; it explicitly leaves out audit and tax work, multi-currency or multi-entity consolidation, external reporting, and how an agent behaves when it needs to ask a clarifying question.

Mercor’s core business is supplying vetted expert labor for AI training and evaluation, so a benchmark demonstrating that expert-built tasks and rubrics are what separates a rigorous eval from a shallow one sits close to its own commercial interest. That does not make the finding wrong, but it is worth naming.

Ramp also shows up elsewhere in today’s issue with its own private engineering benchmark, and two benchmarks from the same company in one day is a signal about where its measurement budget is going.

For a finance team weighing whether to let an agent touch a real ledger, the number that matters is not the 56.4 percent headline score but the 2.6 percent consistency figure beneath it. A model that gets a scripted, graded scenario right on average is not the same as a model that can be trusted to get a real company’s books right every time it runs. Treat APEX-Accounting as a ceiling on current capability, not a certification, and ask any vendor claiming to automate close work for its own repeat-run consistency numbers before granting write access to the books.

Mercor, working with Ramp, published these APEX-Accounting benchmark results in a blog post by Jasmin Kern on July 31, 2026.