By Kevin Boeckholt, CPA, Accordance.
Most of the commentary since the IRS Office of Professional Responsibility released Alert 2026-19, Introductory Guidelines for Responsible AI Use in Federal Tax Practice, on June 24, 2026, has centered on one provision: the treatment of billing under section 10.27(a), which now makes clear that billing a client for manual research time that AI actually replaced can constitute an unconscionable fee.
That’s a real compliance issue, and firms are right to take it seriously. But it isn’t the provision that should be reshaping how firms evaluate AI tools right now. That’s section 10.37, which governs written advice — and which the alert uses to say something firms haven’t fully reckoned with: reliance on an AI system’s output may be unreasonable if that system’s logic can’t be traced to a verifiable source. Opacity, not cost and not billing model, is the standard tax practitioners now need to apply when they evaluate an AI tool. And until that question is resolved, the billing question everyone’s been debating can’t actually be answered.
The alert itself mostly restates existing Circular 230 duties and points them at a new fact pattern: practitioners remain responsible for what AI produces, hallucinated citations can get you sanctioned, client data can’t go to unsecured tools. None of that is new in substance. Due diligence and competence requirements already existed. What’s new is the fact pattern they’re now being applied to, and the opacity problem is where that new fact pattern actually bites.
Why Opacity Is Now a Liability
Most general-purpose AI systems, including the large language models now common in tax and accounting workflows, aren’t built to trace their reasoning back to a verifiable source. They produce fluent, confident answers. They don’t reliably show the specific statute, regulation, or court decision an answer rests on, and when they do cite something, that citation isn’t guaranteed to exist.
The alert cites a real example. A 237-page report prepared for the Australian government by Deloitte Australia, published in July 2025, contained invented quotes attributed to a judge and citations to academic works that didn’t exist, all apparently generated by AI and unreviewed before delivery. The errors weren’t discovered by the firm. They were discovered by an outside academic who checked the citations against the sources and found that some of them didn’t exist at all. Deloitte Australia ultimately refunded part of its fee to the client.
The point for tax practitioners isn’t that Deloitte’s error was unusually careless. It’s that the failure was invisible until someone went looking for the underlying source, and by then it was in a government report. A tax memo or a return position built the same way carries the same risk, just with a smaller audience and, for now, less scrutiny. A firm doesn’t need a 237-page report or a government client to have the same failure mode sitting in a client file.
Two more provisions raise the bar further. Alert 2026-19 applies section 10.22‘s due diligence requirement to mean practitioners must verify the accuracy of facts, citations, and calculations before AI-assisted work reaches a client or the IRS. It applies section 10.35‘s competence requirement to mean practitioners must understand how an AI system generates content and judge whether its outputs are fit for IRS matters. “Understanding how a system generates content” cannot mean reading a vendor’s marketing page or trusting a benchmark score; it has to mean being able to answer, for a specific piece of AI-assisted work, where each factual and legal claim came from and whether that source is still good law. Section 10.22’s verification duty and section 10.35’s competence duty converge on the same practicality: someone at the firm has to independently confirm what the AI told them, using the same primary-source research a preparer would have done by hand.
When the tool itself can’t produce that trail, the firm builds it manually, re-deriving citations, re-checking calculations, cross-referencing regulations the AI already claimed to have relied on. Essentially, the verification workload that the tool was purchased to remove in the first place reappears as a second layer of work sitting on top of the first. A preparer using an opaque tool is, in effect, paying twice for the same research: once to generate it, and once to confirm it actually happened the way the output implies.
What “Traceable” Means in Practice
It’s worth being concrete about what separates a compliant tool from a noncompliant one here. A general-purpose AI system, asked a tax research question, typically returns a fluent paragraph of analysis with a conclusion stated confidently in the first or last sentence. If it cites anything, the citation appears as text: a case name, a Revenue Ruling number, a Code section. Nothing about that output lets a preparer confirm, without leaving the tool and doing the research independently, that the citation exists, that it says what the AI claims it says, or that it hasn’t been superseded. The preparer is asked to trust the citation on the strength of how confidently it was stated, which is precisely the failure mode the alert is warning against.
A system built for traceability instead treats the primary source as the product, not an afterthought attached to one. Every factual or legal claim in the output links to the specific statute, regulation, ruling, or case it rests on, in a form the preparer can open and read directly, in the original language, without taking the AI’s characterization of it on faith. It changes what the section 10.22 verification step actually requires: instead of re-deriving the research from scratch, the preparer confirms a specific, already-surfaced source says what the tool says it says. That’s a materially smaller task, and it’s the kind of task Circular 230 due diligence has always contemplated a competent professional performing quickly, not the open-ended research project an opaque tool leaves behind.
This is a solvable engineering problem, not a promise about model behavior, and it is worth being specific about how. Some systems now separate the component that retrieves and reads primary sources from the component that drafts the answer, so that the draft is composed strictly from a finished research record rather than from the model’s own memory of tax law. Within that structure, every quoted passage can be tagged with a marker tied to the specific document and section it was actually pulled from, and that marker, rather than free text, is the only form of citation the drafting component is permitted to use. At the point the answer is assembled, each marker is checked against what was genuinely retrieved, and one that doesn’t match simply does not appear in the output. That distinction, between an instruction a model can drift away from over the course of a long answer and an architectural constraint it cannot route around, is the practical difference between a tool a firm can supervise under section 10.35 and one it can only hope to.
Why the Billing Question Can’t Be Solved First
This brings the argument back to section 10.27(a), where the commentary has concentrated. The guidance is direct that cost savings from AI should be passed on openly, with billing that reflects the efficiency actually gained, and some practitioners are reading that as a signal that the profession should move toward value-based billing generally. That may be right eventually, and the broader shift away from hourly billing predates this alert. But applying that shift here gets ahead of itself.
A firm cannot credibly price by value delivered, or even document the efficiency it owes a client under section 10.27(a), if the tool that generated the work can’t show how it arrived at its answer in the first place. You cannot bill for what you cannot verify, and you cannot verify what you cannot trace. Opacity isn’t only a section 10.37 problem. It’s the reason the billing question is premature until traceability is solved.
What Firms Should Do Now
Alert 2026-19 puts one question at the center of every AI tool evaluation that wasn’t necessarily there before: can this system trace its answer back to a specific, checkable primary source, or does it just produce a confident paragraph? Firms don’t need to wait for further OPR guidance to start acting on that question.
Any AI tool evaluation or renewal conversation should include a direct question to the vendor: for a given output, can the system show the specific primary source behind each material claim, in a form the preparer can open and check, or does it only produce citations as unlinked text? Firms should also start documenting, at the engagement level, which claims in AI-assisted work were independently verified against a primary source and by whom, since that record is exactly what section 10.22 due diligence and section 10.35 competence will be judged against if the work is ever questioned.
And firms should treat verification time as a real cost of using an opaque tool, accounting for it honestly rather than assuming the advertised time savings are fully realized once the human review layer is added back in.
Everything else the alert touches — verification workload, billing defensibility, exposure under section 10.37 — follows from the answer to that one question. A tool built for tax research, one that shows its work against primary sources by design rather than as an add-on, isn’t just easier to use. It’s the only kind of tool that lets a firm actually meet the standard the IRS has now put in writing.
===
Bio: Kevin Boeckholt, CPA, leads Accordance‘s innovation practice, where he works with tax and accounting professionals to apply AI to the day-to-day realities of their work. He previously led product at an early-stage tax AI company and served in Tax Technology Consulting at Deloitte.
Sign in to get access to this free resource, and all of our whitepapers and reports.
Download this content today!
Register Now Already registered? Click here to Log In