A year ago, artificial intelligence tools for SOX testing could barely handle a clean three-way match. Since the end-of-2025 model releases, they have been working through complex reconciliations and hundred-megabyte spreadsheets, and the Big Four are building their own: KPMG has published a framework for generative AI in SOX, Grant Thornton a seven-step path, and Deloitte and EY each have multibillion-dollar AI audit platforms in progress. Every vendor now promises agents that read the evidence, test the samples, and hand a finished workpaper to a reviewer.
The buyers are under pressure to say yes. KPMG’s 2025 SOX Survey puts the average program at $2.3 million and more than 15,000 hours a year (~40% of them on testing), and the Internal Audit Foundation found 57% of internal audit functions holding headcount flat while their mandate grows.
The trouble is that the demos all look alike, and you learn which kind of tool you bought months later, when the workpapers are due. In my book, The Audit Leader’s Guide to AI for SOX Testing, I boiled the evaluation down to five questions, all of which you can ask in the first vendor meeting.
1. Is it an agent, or a chatbot with a nicer interface?
A chatbot answers a question. An agent is given a goal, the rules for reaching it, and a set of tools, and it returns structured work a reviewer can inspect. A SOX test is a chain of steps, and a long prompt can describe the chain but cannot reliably run it.
Many teams tried prompt libraries and got stuck: nobody owned the prompts, nobody could measure the output, and every control needed its own. Worse, a model that is right 95% of the time is useless if you cannot tell which 5% to check. Then you check everything and redo the work the AI was meant to save.
2. Can it handle real-world evidence, and does it know where to start?
SOX evidence is a mess by nature: screenshots, PDFs, spreadsheets, emails, and system exports with no naming convention, and about one in five PBC requests coming back wrong the first time. This is what broke robotic process automation, which Deloitte found only 3% of companies ever scaled.
So bring the ugliest evidence from your own program to the demo, and turn down the vendor’s sample files. Then see yourself where the tool works today.
3. Does the audit trail have all five layers?
AI is not fully predictable. Run the same test twice, and you can get slightly different answers. A tool worth trusting is built with that in mind, and the design has five layers:
- An approved testing plan: the tool writes out the procedures, and an auditor approves them before testing starts.
- The source evidence, stored exactly as received.
- An extraction layer that can answer four questions about every fact it used: what was extracted, from which file, where in that file, and how.
- An assessment layer where those facts become a result with the reasoning in plain English.
- Human review, which is where most of the auditor’s time now goes.
Ask to see all five on one of your own controls.
4. Will the external auditor accept it?
There is still no PCAOB standard for AI in audits, and the staff’s July 2024 observations never became guidance. That leaves AS 1215 as the bar: the file has to be detailed enough that an experienced auditor who has never seen the engagement can follow what was done and why. A tool built on the five layers passes this bar.
Bring the external auditor team early. Lead with the fact that your auditor still signs off. Run one parallel pilot and let them compare the two files. Start with controls they already know well. And keep in mind that every major firm is building AI testing tools of its own: they believe the technology works, but they will look at how you apply it.
5. What does a fair 90-day pilot look like?
Don’t pick a tool based on the best demo, and don’t start with the hardest controls. Choose five to 10 controls and one person to own the pilot. Record your baseline (hours per control, PBC follow-up rate, rework hours, time to a final workpaper) before committing to a vendor. Then, run the old way and the new way on the same controls and compare time, documentation quality, how the reviewer felt about signing off, and whether both runs caught the same exceptions.
Measure testing quality, hours given back to the team, and fewer repeat PBC requests. A well-chosen vendor lets you spend more time on deep work, and pays back in four to six months.
The one standard behind all five
The five questions reduce to a test the profession already applies to a new hire or a co-source firm: show me the plan you followed, the evidence you used, the reasoning behind each conclusion, and the reviewer who signed off. A tool that can do that on one of your own controls has earned a pilot. One that can has not, no matter how good the demo.

ABOUT THE AUTHOR:
Alexey Zanin is the CEO and co-founder of Bead AI and the author of The Audit Leader’s Guide to AI for SOX Testing, a free, vendor-neutral guide for audit leaders. He previously led regulatory compliance programs at Meta.
Photo credit: TSD Studio/Unsplash
Sign in to get access to this free resource, and all of our whitepapers and reports.
Download this content today!
Register Now Already registered? Click here to Log In
Tags: ai agents, Artificial Intelligence, audit, Auditing, SOX, Technology