AI Accounting Work

The AI Scored 100%. The CPAs Scored 37%. Read the Appendix.

Mercor published a study in October 2026 that produced an irresistible headline: AI now outperforms accountants at month-end close. Twelve US accountants — every one holding an active CPA license, half with Big Four experience — attempted four simulated close scenarios: nonprofit budget variances with shared-cost allocations, a hotel revenue and commission reconciliation with correcting entries, an event revenue and business-mix analysis, and lease rollforwards with journal entries. Across 23 unassisted attempts, the humans averaged 37% of rubric criteria met. Claude Opus 5 scored 100% on all 20 of its attempts, finishing each in under ten minutes, at $0.21 per correct rubric criterion versus $10.35 for the humans.

"It's over for accountants" trended within hours. The paper is more interesting than the headline — in both directions. I read all fifteen pages, including the appendix, which is where the real story lives.

What the 37% actually measures

Apply the first question from the field guide for auditing dramatic claims: is the measurement what the headline says it is?

Here, the measurement is cold-start performance on a stranger's books. Participants got a simulated company's working files, no coworkers to ask, no accumulated context, and a rubric where errors cascade — the tasks were adapted from a benchmark originally designed to stump AI models, dense with hard-to-spot traps, so one missed upstream adjustment sinks every downstream answer. Scoring used an LLM judge (spot-checked by hand on eight submissions, with 98% agreement — reassuring, though not a full audit).

And here's the number that reframes everything: the task authors predicted humans would score about 30% for juniors and 50–60% for mid-levels. The participants averaged 37%. The humans didn't fail the benchmark — they performed almost exactly to spec. The paper's own authors say it plainly: the tasks tested "precisely what the models are best at" — being detail-oriented, searching the whole file set, following instructions exactly — under conditions that strip away everything humans lean on. One participant put it better than any methodology section: "I am in month-end close right now and I could not imagine getting these tasks done in less than 3 hours."

So the story isn't that CPAs are bad at their jobs. The story is the other line on the chart: eighteen months before this study, the best AI models scored below the humans' 37% on these same tasks. Now they ace them. The humans stood still — the models moved.

One more appendix detail worth knowing: the paper's title says "junior accountants," but Appendix A2 lists 58% senior accountants or senior auditors and 25% managers or above — only two of twelve were staff level. The result is stronger than the title admits.

The part the profession shouldn't wave away

The comforting rebuttals — cold start! compounding rubrics! no coworkers! — are all true, and all beside the main point. A task author quoted in the paper closes that exit: "The errors themselves are not rare. Something coded to the wrong account, a bill that lands after the month ends, a commission that never made it into the books — that is a normal close, and an accountant sees them constantly." Seven of twelve participants rated the tasks realistic, and the two who said "very unrealistic" clarified they meant the setting, not the work.

Reconciliations, rollforwards, correcting entries, variance analysis: this is not edge-case accounting. It's the bulk of staff and senior preparation hours in any close calendar. The study didn't test the whole job — but the part it tested is the part that fills the timesheet. Even restricted to the rubric items humans tend to get right, the models were 34 times faster. "Someone still has to sign off" remains true, and remains irrelevant to headcount: sign-off at today's staffing levels assumes humans do the preparation underneath the signature.

The buried lede: humans plus AI lost to AI alone

The study was designed to measure augmentation — how much AI helps a professional. Instead it found that accountants working with Claude were about 15× slower than Claude alone, and slightly less accurate.

The detail that explains it is the most important paragraph in the paper. Of the four AI-assisted sessions that missed a perfect score, in three of them Claude had the correct figure at some point and it didn't survive to the final submission. Twice, Claude revised away from its own right answer and the human deferred to the new wrong one. Once, the human overrode Claude's correct advice directly.

Both failure modes are the same failure: review without evidence. A reviewer who can't trace a number to its source document can only add noise — second-guessing good work, rubber-stamping bad revisions, deferring to confidence instead of support. That's not an argument for removing human review. It's an argument that the only review worth doing is evidence-based: every AI-prepared figure carries a link to the document it came from, and the human checks exceptions against sources rather than re-deriving the close or trusting the vibe.

For anyone building these tools, the architecture follows directly: the AI proposes, a deterministic system posts, and every entry carries its provenance — which source, which rule, which approval. Autonomous posting because twenty benchmark runs went perfectly is how restated quarters happen. Source-linked preparation with exception-driven human approval is how the ten-minute close actually gets banked.

What shifts, honestly

The work that compresses: manual preparation. The recs, the rollforwards, the first drafts of entries — the model does them faster, cheaper, and per this study more meticulously than practicing CPAs under benchmark conditions. That compression is real, and the model-progress curve says it arrived in about eighteen months.

The work that compounds: everything the study explicitly could not measure. The paper's own list — client communication, asking the right questions, operating under ambiguity and missing evidence, accumulated context — plus controls design, exception judgment, and owning the number when the auditor or the bank asks who stands behind it. The signature was never the hard part of sign-off; the accountability was.

The professionals who get hurt are the ones whose entire value is the hours the model just did in ten minutes. The ones who do fine will treat the model like a brilliant, tireless, occasionally overconfident preparer: trust the work, verify the sources, and never let it post unsupervised.

Sources: Mercor's study summary · full paper (PDF)