Frontier AI models completed four simplified month-end accounting scenarios faster and more accurately than 12 licensed accountants, Mercor reported in a study published October 1, 2026. The models scored perfectly. But that striking result measures a narrow slice of accounting work—not the client conversations, teamwork and accumulated company knowledge that the study left out.
Perfect scores changed the experiment
Mercor recruited certified public accountants, or CPAs, averaging about five and a half years of experience. Their assignments came from simplified versions of its APEX-Accounting benchmark. Each required searching a company’s working files, finding the right figures, doing calculations and delivering a table of results.
The researchers originally wanted to measure whether AI improved human performance. Because models alone achieved perfect scores, Mercor said there was no room to measure that improvement. The published comparison instead focuses on unassisted humans versus AI; it does not establish the gains from a combined accountant-and-AI workflow.
Mercor also found a large cost advantage: frontier models were more than ten times cheaper when comparing cost per grading criterion met. Its historical comparison shows how quickly performance changed. Eighteen months earlier, the best models fell below the accountants’ average score of about 37%.
the traps [hidden in the task] are realistic, and the conditions and the scoring are what drove the results.
An accounting task author, quoted by Mercor
Realistic traps, stripped-down working conditions
The assignments contained hard-to-spot accounting requirements that compounded one another. A single missed number could produce a very low score. Task authors expected average scores of just 30% for junior accountants and 55% for mid-level accountants—not near-perfect human results.
Accountants had no coworkers to ask for help and no accumulated company context. Mercor says those conditions made the tasks harder. The test also excluded several responsibilities that could become more important as AI handles structured work:
- Communicating with clients.
- Asking the right questions.
- Building context and knowledge that are not written into the task instructions.
Mercor identifies a tension in how these tests evolve. Benchmark designers make tasks harder to expose model failures; some now require multiple experts and tens of hours to build. One person may need comparable time and help to finish them. The company argues that frontier benchmarks are consequently shifting toward work previously impossible or prohibitively costly, rather than simply measuring how well AI performs existing human jobs.
The larger benchmark remains unfinished
The full APEX-Accounting benchmark contains 160 tasks across 10 simulated companies. Its scope includes reconciling accounts, accruing expenses, posting transactions and producing reports, according to the benchmark paper.
Each simulated company has an accounting system, spreadsheets, PDFs and other files. Accounting and bookkeeping experts both authored and solved every task, then wrote the grading rules. That setup tests work across a collection of business records, not just answers to standalone accounting questions.
In its October 2 coverage, The Decoder reported that Claude Opus 5.5 led that benchmark with 61.8% of grading criteria met. Fable 5.1 followed at 61.0%, with GPT-6 Astra at 57.9%. It also reported Mercor’s finding that almost 60% of tasks had not been fully solved by any model. Meeting some grading criteria is different from completing an assignment.
The earlier July benchmark paper also tested repeatability: its best result for completing a task correctly in all eight runs was only 2.6%. The highest score for completing tasks in at least one of eight attempts was 21.5%, achieved by a different model.
Those are historical results, not October model scores. They illustrate another distinction the simplified study cannot settle: strong performance on selected tasks versus dependable completion across repeated attempts.
Reader comments
Newest comments first. Replies stay oldest first.