The Decoder is reporting a Mercor study. Twelve licensed CPAs, with an average of five and a half years of experience, worked through simplified tasks from the APEX accounting benchmark. Eighteen months ago, the best models still scored below the accountants’ average of about 37 percent. The article then says that today the models solve those same tasks almost flawlessly. That sentence is about the simplified tasks, not the full benchmark below. The article also says that on structured bookkeeping the models are faster, more accurate, and far cheaper than accountants. No dollar figure is given for the cost, so none is added here. A comparison chart is an image. The article text does not list a score for each model, and those scores are not filled in.

[1]
A fully dark booklet with a dull red edge, beside a taller stack that is marked only across the top and blank below.
The booklet on the left is dark all the way down. The taller stack is marked only at the top, and the rest is blank. That is the simplified set and the full set as two different results. An illustration, not a table., AI-generated illustration, not a news photograph

The full APEX accounting benchmark is much larger: 160 tasks across 10 simulated companies, built by more than 40 professionals who average 11 years of experience. Claude Opus 5.5 currently leads, meeting 61.8 percent of the grading criteria, followed by Fable 5.1 at 61.0 percent and GPT-6 Astra at 57.9 percent. Mercor says no model fully solved almost 60 percent of the tasks. The 61.8 percent is a share of grading criteria met. The almost 60 percent is a share of tasks not fully solved. The two figures use different units. They are not converted into each other, and “almost flawlessly” is not moved onto the full benchmark.

[1]

Mercor admits that the tasks in the study are what these systems do best: finding details and following instructions precisely. Left out are talking with clients, checking with colleagues, and using context built up over years. Mercor says that is why accountants cannot be replaced. It also expects productivity gains across the industry. The article does not say how large those gains would be.

This is The Decoder’s account of the Mercor study and the APEX leaderboard, not the original paper or the leaderboard page.

[1]

要点

  • On the simplified tasks, the best models were below the accountants’ roughly 37 percent eighteen months ago. The article says today’s models solve those same tasks almost flawlessly.
  • Full APEX has 160 tasks. Opus 5.5 meets 61.8 percent of grading criteria, Fable 5.1 meets 61.0 percent, and Astra meets 57.9 percent.
  • Mercor says no model fully solved almost 60 percent of the tasks. That is not the same result as “almost flawlessly.”
  • The tasks leave out clients, colleagues, and years of context. No cost figure is given.