AI Beats 12 Licensed CPAs on Bookkeeping Speed and Accuracy
Think about the last time you sat down with a pile of receipts and a spreadsheet. Matching numbers, tracking down a missing invoice, double-checking that the totals actually add up — that's the part of accounting nobody enjoys. According to a new study from Mercor, it's also the part where AI for accounting has quietly raced ahead of humans. AI models now beat 12 licensed accountants on speed and accuracy at structured bookkeeping tasks, and they're far cheaper while they do it. But before you fire your bookkeeper, read the fine print: none of these models can close the books without someone watching over them.
What Happened: From Below 37% to Almost Flawless
Mercor ran a study with 12 licensed CPAs — certified public accountants, the people who are legally allowed to sign off on financial statements. On average, they had about five and a half years of experience each. Every one of them worked through simplified tasks taken from something called the APEX Accounting Benchmark.
Think of that benchmark as a driving test for AI. It hands a model a realistic accounting job — reconcile this account, sort this transaction — and then grades the answer against a checklist a human reviewer would use. It's not about whether the answer looks pretty. It's about whether it's right.
Eighteen months ago, the best AI models scored below the accountants' average of about 37 percent. They lost, and it wasn't close. Today, those same models solve the same tasks almost flawlessly.
The Full Benchmark Tells a Messier Story
Here's where it gets more interesting. Those simplified tasks aren't the whole exam. The full APEX Accounting Benchmark has 160 tasks spread across 10 simulated companies, and it was built by more than 40 professionals who average 11 years of experience each.
On that harder version, no model has come close to acing it:
• Claude Opus 5.5 leads with 61.8 percent of grading criteria met.
• Fable 5.1 is right behind at 61.0 percent.
• GPT-6 Astra sits at 57.9 percent.
And according to Mercor, there are almost 60 percent of tasks that no model fully solved. So the same technology that looks perfect on small, clean jobs still stumbles hard on big, messy ones.
What the Study Left Out
Mercor is refreshingly honest about this part. The tasks it used test exactly what AI is best at: hunting down details and following instructions to the letter. What they don't test is the rest of the job — talking with clients, checking in with colleagues, and drawing on context someone built up over years of working with the same company.
That's the human part. It's also why Mercor says accountants aren't getting replaced, even though it expects big productivity gains across the industry.
What It Means for You
You don't need to run a finance department for any of this to matter. Here's how it lands in real life.
At home. If you've ever tried to sort a year of receipts before tax season, a model can categorize transactions in minutes — as long as you feed it a clean spreadsheet. You still have to eyeball the weird ones.
At work. If part of your job is coding invoices or matching payments, your day is about to change shape. The mechanical first pass gets faster, and your judgment on the exceptions becomes the actual job.
For your business. A small business owner can now get a first-pass bookkeeping review without waiting on anyone's calendar. There's a free way to start below — just don't treat the output as final.
For students. Accounting students finally have a tireless study partner. Feed it practice problems, ask it to explain why an entry is wrong, then verify against your textbook, because it will occasionally sound confident and still be wrong.
For creators. Freelancers and creators live and die by invoices and expense tracking. Sorting a year of royalty statements and platform payouts turns into a copy-paste job instead of a lost weekend.
For your income. Bookkeepers and accountants who learn to run these tools can take on more clients without hiring anyone. The ones who ignore them will get compared to a machine that costs almost nothing per task.
If you want to poke at this without spending a dime, mykreatool.com has a stack of free AI tools worth trying before you commit to anything paid.
How to Try It Right Now
You don't need an accounting degree. You need a clean file and a habit of double-checking.
1. Start free. Head to mykreatool.com and pick a free tool for a small test run. Free is the right place to start — you're testing the workflow, not the ceiling.
2. Pick one small job. Export 20–30 transactions from your bank or accounting software as a CSV. Small and boring beats big and complicated for a first attempt.
3. Write a specific prompt. Something like: "Categorize each row into one of these categories, and flag anything you're unsure about with a question mark." Vague prompts get vague answers.
4. Ask for confidence flags. Tell it to mark the rows it isn't sure about. You want the exceptions, not a tidy-looking list that's quietly wrong.
5. Check the flagged rows yourself. This is the step people skip, and it's the one that matters. A model that's right 95 percent of the time is still wrong somewhere in your money.
6. Move up only if it pays off. If the free run saves you real time, the models at the top of the APEX leaderboard are Claude Opus 5.5, Fable 5.1 and GPT-6 Astra. Test them on your own data before trusting any of them with a full set of books.
Upsides and What Changes
The upside isn't "AI replaces accountants." It's that the boring half of the work gets cheap. Detail-hunting and instruction-following are exactly what these models do well, and those skills eat up an enormous chunk of an accountant's week.
What actually changes:
• First-pass work gets fast. Sorting, categorizing and cross-checking stop being the bottleneck.
• Cost per task drops hard. Mercor's finding is blunt: models are far cheaper than humans on structured tasks.
• Review becomes the real job. Value shifts from doing the work to knowing which answers to trust.
• Small operations get bigger leverage. A one-person business can now handle bookkeeping that used to eat an hour of a professional's time.
None of that is hypothetical. It's what happens when benchmark scores jump from below 37 percent to nearly perfect in 18 months.
Limitations
Let's be honest about what this study is and isn't. The tasks were simplified, and Mercor itself says they test precisely what AI does best — detail work and rule-following. The benchmark left out client conversations, colleague coordination, and the years of context a good accountant carries around in their head. On the full 160-task benchmark across 10 simulated companies, the best model met 61.8 percent of grading criteria, and almost 60 percent of tasks weren't fully solved by any model. That's a genuinely useful assistant, not a replacement. Anyone claiming AI can close your books unsupervised is selling something.
Conclusion
AI has honestly overtaken licensed accountants on speed and accuracy for structured bookkeeping — from below 37 percent 18 months ago to nearly flawless today. But "nearly flawless at the easy tasks" is a long way from "runs your finances." The win right now goes to the person who uses the machine for the first pass and saves their own judgment for the exceptions.
Your one action today: grab 20 transactions from last month, run them through a free tool at mykreatool.com, and count how many rows you have to fix yourself. That number tells you more about your next 12 months than any headline will.



Comments 0