- feature
- PRACTICE & PROCEDURES
A risk framework for AI use in tax administration and preparation
A former IRS commissioner assesses risks and proposes control practices for use of artificial intelligence by taxing authorities and practitioners.
Related
Rev. Proc. 2026-17: Withdrawal of Sec. 163(j)(7) elections
Tax legislation for digital assets: What’s the conversation in Congress?
Tax return positions: Required levels of authority
Artificial intelligence (AI) has the potential to disrupt virtually every industry sector and business function. The opportunities for improvements to organizational performance in areas such as productivity and customer satisfaction could be massive and unprecedented. However, these opportunities come with an array of risks that are not yet fully understood. As the developers of AI products and the users of those products race to unlock potential performance gains, there is an emerging need for domain–specific risk frameworks in areas such as legal, medical, insurance, accounting, tax, and financial services to guide organizations as they move from AI strategy to deployment.
Of the various business areas that are early movers in AI, tax may be the most AI–compatible professional domain. The reasons are structural: Tax runs on deterministic input–output logic; sits on hundreds of millions of labeled interactions; is anchored to a body of authoritative content in the Internal Revenue Code, Treasury regulations, and state tax laws; and operates at a scale that drives substantial projected results in AI business cases.
But failure modes in tax are high stakes. Misalignment between AI deployments and the underlying risk associated with them can result in both legal and financial harm at scale. I have experienced this tension while leading a tax authority and working closely with tax preparers. In my numerous discussions over the past several years across the tax community concerning AI, a common sentiment persists, voiced something like: “We understand that AI is a game changer. We have seen evidence that it can dramatically improve performance across practically all key performance areas. But we don’t know all the risks. We don’t know what questions to ask to assess our own readiness for responsible AI adoption, nor do we know what to ask the AI vendors as part of our due diligence in ensuring their product and services will be deployed responsibly.”
As a result, it has been consistently suggested to me that it would benefit the tax community to have a common risk framework to work from — a starting point where best practices can accumulate, learnings can be shared, and the framework can be improved collaboratively as the AI and tax environment evolves.
This article provides a framework to address these needs. It establishes a notional risk register for tax authorities and presents that side by side with a distinct register for tax preparers. It also identifies important areas of overlap across the two. The main text of the article is intended to introduce these risk frameworks, while the appendices contain detailed methodologies and tools for those interested in applying the frameworks to specific AI deployment decisions.
Risk registers for the use of AI in tax
A risk register is more than a list of relevant risks. The formulation and activation of the list means the organization has learned to:
- Define how risks show up in practice;
- Assign controls;
- Establish indicators; and
- Make course corrections as the environment evolves.
In this way, the register is part of a broader risk governance framework to help organizations chart a sensible path forward from strategy to implementation. If effective, the proposed register and associated framework will assist tax administrators and practitioners in identifying the right set of questions to conduct both(1) a self–assessment of their readiness for responsible AI adoption and (2) a due–diligence assessment of the products being marketed to them by AI vendors.
To this end, two distinct frameworks are offered — one for tax authorities and one for tax preparers.
AI risk register for tax authorities
The Tax Authorities risk register presented here includes 20 risk areas organized into the following four categories:
- Information integrity: The wrong information reaching taxpayers at scale.
- Fairness and legitimacy: Disparate or unfair treatment across groups.
- Security and data: Compromised systems, corrupted data, or the exploitation of sensitive taxpayer information that AI systems aggregate, surface, or make newly accessible to bad actors.
- Institutional capacity: Loss of the knowledge and capacity by the institution needed to govern AI responsibly over time.
Of note, the risks presented in the register address the authority’s own AI deployments. However, tax authorities must also maintain a defensive posture against AI being inappropriately used by taxpayers when filing. This includes well–meaning taxpayers acting on bad or poorly designed AI advice and bad actors using AI to commit fraud. Several of the deployments highlighted in the register — including audit selection tools, pattern detection, and machine–learning–driven screening — are themselves the authority’s primary controls for detecting and responding to AI–influenced tax filing at scale. A comprehensive treatment of AI risk for tax authorities may ultimately warrant the creation of a separate register focused on these external threats.
Figure 1, “The 20 Risks in the Tax Authorities Risk Register Defined,”presents each risk defined by describing the failure mode to be avoided.
Figure 1: The 20 risks in the Tax Authorities risk register defined1
I. Information integrity
1. Hallucination: The AI generates confident, specific, and factually incorrect output, whether delivered to a taxpayer seeking guidance or embedded in an authority’s internal workflow.
2. Drift: The AI’s outputs gradually shift away from accurate tax interpretation over time, without any single failure triggering a review.
3. Explainability: The AI cannot trace or explain why it produced a given output, making it impossible to verify, audit, or defend.
4. Over-trust/unauthorized reliance: Staff, taxpayers, or preparers treat AI output as authoritative, removing the human check that catches errors — and in the most consequential cases, creating potential estoppel claims.
5. Policy drift: The AI subtly reshapes how tax rules are interpreted through accumulated outputs, effectively rewriting policy without any formal change process.
II. Fairness and legitimacy
6. Bias/neutrality:The AI produces systematically different outcomes for different taxpayer groups based on income, language, geography, or other characteristics unrelated to legal merit.
7. Audit targeting: The AI introduces or amplifies bias in how taxpayers are selected for audit, producing disparate impact across industries, demographics, or geographic groups.
8. Trust shock/compliance norm erosion: A visible AI failure triggers rapid loss of public confidence in the tax authority — and accumulated smaller failures gradually erode voluntary compliance.
9. Surveillance perception: Expanded AI use creates public perception that the authority is conducting mass surveillance, eroding trust even where the underlying use is legitimate.
Figure 1: The 20 risks in the Tax Authorities risk register defined (continued)
III. Security and data
10. Prompt injection: A bad actor manipulates inputs to the AI system to bypass safeguards, extract sensitive info, or produce harmful outputs.
11. Data leakage: Taxpayer data entered into or processed by AI systems is exposed, accessed without authorization, or retained and reused by vendors in ways that violate privacy obligations — and the aggregation of that data creates concentrated breach targets.
12. Re-identification: Anonymized data processed by AI systems is combined with other information to identify specific individuals, creating unanticipated privacy exposure — and AI accelerates the speed and scale at which re-identified data is weaponized for fraud.
13. Adversarial manipulation: Sophisticated actors exploit known weaknesses in AI detection models to evade compliance enforcement.
14. Agentic AI risk: An AI system with autonomous action capabilities takes irreversible actions before a human can intervene, is manipulated into accessing systems or data beyond its intended scope, or causes cascading failures across connected legacy systems at a speed that outpaces human oversight.
IV. Institutional capacity
15. Workforce and expertise erosion: AI adoption reduces hiring and development of human tax expertise while senior institutional knowledge departs, leaving the authority less capable of governing the systems it has deployed.
16. Scope creep: AI systems are quietly expanded beyond their originally approved use cases without corresponding updates to controls, governance, or validation.
17. Control dilution: Existing review/oversight processes fail to scale with AI-driven increases in volume and speed, resulting in weaker governance even if all controls remain unchanged.
18. Lack of auditability: The authority cannot reconstruct what its AI systems did, making appeal response and error identification implausible.
19. Regulatory and policy noncompliance: The authority deploys or operates AI in ways that are inconsistent with applicable laws, regulations, and policies.
20. Enterprise governance gap/gradual autonomy creep: No single function within the authority is responsible for managing AI investments across the organization — and AI tools gain autonomous capabilities incrementally without triggering a formal governance review.
The complete register is presented in Appendix A and includes, for each risk, a description, an illustrative tax scenario, and illustrative controls.2
The proposed register is based on a literature review of both AI– and tax–related risk; a review of established risk frameworks (e.g., the COSO Enterprise Risk Management (ERM) model);3 work undertaken by the IBM Center for the Business of Government;4numerous and ongoing discussions with tax authority personnel; and direct experience leading a large tax authority. It is worth noting that “AI” encompasses a wide range of tools — generative models, machine learning, natural–language processing, and others — each with distinct risks. Responsible deployment begins with choosing the simplest method suited to the task.
A list of potential risks is far more useful when connected to a framework that guides its use in practice. Presenting the risks in a quadrant view in Figure 1 was by design. This grouping reinforces that, in the public sector, risks often present simultaneously, are interdependent, and can be in tension with one another. Specifically, mitigating risk in one area can increase it in another. For example, making AI use more transparent can lessen “explainability” risk but increase “adversarial manipulation” risk, as disclosing how detection models work could enable bad actors to evade them. Similarly, greater transparency could increase “surveillance perception” risk. Thus, the proposed quadrant view of the risk register is grounded in the COSO ERM framework’s treatment of risk as multidimensional rather than sequential.
A limit of the quadrant view is that it implies that all risks carry equal weight and should be given equal attention. To address this limit, the proposed framework includes a three–tiered system where Tier 1 risks must be addressed before a deployment launches, and Tier 2 and 3 risks are mitigated during operations but do not typically affect “go/no–go” decisions. More detail on the tiering process is in Appendix B. Figure 2, “The Full AI Risk Landscape for Tax Authorities,” shows how the risk register may look after risk tiers are assigned.
Figure 2. The full AI risk landscape for tax authorities
20 risks organized across four quadrants, color-coded by risk tier.

With a priority structure in place, the risk register can now be plugged into a broader set of steps that yield the optimal set of questions the organization should ask and answer when deploying an AI solution. A suggested eight–step methodology is described in full in Appendix B and applied to a specific case study. The essential components of this methodology are presented more generally here:
- Determine scope: Define distinct boundaries of the AI product. A chatbot authorized only for filing deadlines presents a different risk profile than one authorized to answer credit eligibility questions. Thus, scope definition is essential for determining which risks will apply.
- Determine which risks apply: With scope defined, the tax authority works through the register to identify risks applicable to the deployment, as not all risks are relevant to every solution. A chatbot authorized only to handle filing deadlines would not trigger the “unfair audit targeting” risk. This risk, however, would be triggered by a product that assists in audit case selection.
- Prioritize the risks: Prioritization is an essential element of effective risk management. A public-facing chatbot answering credit eligibility questions would likely result in a tax authority flagging “hallucination” as a high-priority risk. The likelihood of incorrect outputs is material, the harm to taxpayers may be irreversible, and there is potential legal exposure for the tax authority. By contrast, “trust shock/compliance norm erosion” may still be flagged as relevant for the same deployment but would score lower because the impact manifests more gradually.
- Assess organizational governance and vendor accountability: With risks identified and prioritized, the tax authority now must determine the appropriate mitigation controls to put in place. The higher priority the risk, the more intense the controls. The controls fall into categories such as human oversight, review frequency, and deployment phasing. For a chatbot deployment and an associated risk of “over-trust/unauthorized reliance,” controls might include clear legal disclaimers on every output, an escalation path for scenario-specific questions that approach the threshold of legal advice, and a phased rollout that restricts the chatbot to low-stakes factual queries in Phase 1 before expanding to credit eligibility guidance only after Phase 1 accuracy thresholds are verified.
AI risk register for tax preparers
The proposed Tax Preparers risk register includes 15 risk areas organized into the following five categories.
- Technical:Incorrect or unreliable AI outputs that establish the conditions for downstream errors.
- Practice: Failures of professional judgment, documentation, or oversight.
- Legal and regulatory:Increased exposure to legal liability (e.g., malpractice claims, regulatory sanctions).
- Business:Erosion of client trust and firm relationships.
- Workforce and strategic:Loss of the expertise and institutional knowledge needed to practice responsibly.
Presented in Figure 3, “The 15 Risks in the Tax Preparers Risk Register Defined,” is each risk area defined by describing the failure mode to be avoided. The complete register is presented in Appendix C and includes, for each risk, a description, an illustrative tax scenario, and illustrative controls.
Figure 3. The 15 risks in the Tax Preparers risk register defined5
Technical
- Hallucinated advice: The AI suggests a tax position, deduction, or interpretation that is confidently stated but factually wrong or legally unsupported.
- Client data leakage: Sensitive client information — income, identity, business details — is exposed to or retained by the AI vendor through prompts, logs, or training pipelines in ways the client never consented to and the preparer may be unaware of.
- Incorrect document extraction: The AI misreads, misclassifies, or omits information from source documents — Forms W-2 or 1099 or Schedules K-1 — producing errors that flow silently into the return without triggering any obvious flag.
Practice
- Failure of professional judgment: The preparer accepts an AI output without independent verification, over-trusting the confident tone of the response rather than checking the underlying authority.
- Misplaced authority/over-trust: The preparer or client attributes more legal weight to AI output than it warrants — treating a probabilistic suggestion as a definitive answer and failing to recognize the limits of what the tool can reliably determine.
- Control dilution: The speed and volume AI introduces causes existing review and sign-off processes to become less rigorous in practice, even if they remain unchanged on paper. More returns get processed with less scrutiny per return.
Legal and regulatory
- Malpractice exposure: A filing error traceable to AI output results in penalties, interest, or audit findings, and the preparer faces professional liability for a return they signed but did not fully verify.
- Liability ambiguity: When an AI-assisted position is challenged, the allocation of responsibility between the signing preparer and the AI vendor is legally unclear.
- Regulatory noncompliance: The firm’s use of AI violates professional standards, licensing rules, or engagement obligations — including unauthorized practice by unlicensed staff relying on AI for legal interpretations.
Business
- Client trust erosion: The client discovers AI was used in preparing their return — particularly after a disallowed position or error — without prior disclosure, damaging the relationship and potentially triggering a dispute over fees or liability.
- Reputational harm: A pattern of AI-assisted errors or a single high-profile failure becomes known in the market, damaging the firm’s standing with existing and prospective clients beyond the individual affected engagement.
- Lack of transparency: The firm has no clear policy on how AI is used, disclosed to clients, or documented internally, leaving clients uninformed and the firm exposed if AI use is later contested.
Workforce and strategic
- Skill erosion: Staff develop expertise in operating AI tools rather than in the underlying tax law those tools are applying. The firm loses the independent analytical capability needed to catch AI errors or handle complex matters without AI assistance.
- Vendor dependence: The firm becomes operationally reliant on a single AI vendor to the point where a vendor outage, pricing change, or contract termination would materially disrupt client service. The firm no longer has the internal capability to operate without the tool.
- Institutional knowledge loss: Experienced preparers retire or leave earlier than they otherwise would as AI reduces demand for their expertise. The knowledge they carry — about complex clients, “edge” cases, and professional judgment — leaves with them and is not replaced.
The proposed register draws on the same literature review and established frameworks that informed the Tax Authorities register, adapted to reflect the distinct professional and legal environment in which tax preparers operate.
Like the Tax Authorities register, the Tax Preparers register is more useful when connected to a practical framework. The preparer’s situation is sufficiently distinct from a government administrator to warrant a separate register.
The risks faced by a preparer tend to be more sequential, where a single AI error in an earlier segment of the preparer process can cascade through the value chain, later affecting professional judgment, client outcomes, and firm liability. For example, an AI system that hallucinates a deduction creates a technical risk at the start of the process chain. But that technical risk propagates upward or forward only if the preparer fails to catch it. It’s when a junior associate accepts the AI suggestion without independent verification that a technical failure becomes a practice failure (i.e., a failure of professional judgment). That practice failure, if the return is filed, produces potential legal exposure — and so on up the chain, causing increasing harm if it remains unmitigated.
This sequential view of risk — where failure in an early layer creates conditions for failure in the next — led to the proposal of a stack model for the Tax Preparers register. The five risk categories in the register appear as five ascending layers as presented in Figure 4, “The AI Risk Stack for Tax Preparers.” The sequencing facilitates risk prioritization: Preparers are incentivized to invest more heavily in controls at the bottom layers. Mitigating risks in the Technical Risks or Practice Risks layer successfully is more effective than mitigations applied when the error moves to the Legal and Regulatory Risks, Business Risks, and Workforce and Strategic Risks layers.
Figure 4. The AI risk stack for tax preparers

With a recognition of how risks escalate in the context of the tax preparation process, preparers can now plug the risk register into a broader set of steps that yield the optimal set of questions the organization should ask and answer when deploying an AI system. A suggested eight–step methodology is described in full in Appendix D and is applied to a specific case study. The essential components of this methodology are presented more generally here:
- Determine scope: Define distinct boundaries of the AI product. An “AI tool for tax preparation” is not a sufficient deployment definition. A tool authorized only to summarize client documents and surface relevant Code sections presents a different risk profile than one authorized to suggest deductions and draft tax positions for incorporation into filed returns.
- Determine which risks apply: With scope defined, the preparer now works through each line item in the register to identify the subset of risks that apply to the deployment being considered, as not all risks in the register may apply to every potential solution. A tool restricted to document summarization would not trigger the malpractice exposure risk at meaningful intensity — there is no filed position to be wrong. That risk is activated when the tool’s output is incorporated into a return that a client relies upon and a taxing authority later examines. Similarly, regulatory noncompliance risk becomes live when junior staff use AI to make interpretive legal judgments without oversight but not for a research tool used under supervision.
- Prioritize the risks: Prioritization of risks is an essential element of effective risk management, and in the preparer context, it should begin at the bottom of the risk stack. A firm deploying an AI tool to draft tax positions would spend considerable time assessing risks in the technical layer: If the model hallucinates a citation, misreads a statutory threshold, or applies an outdated regulatory interpretation, it will result in significant downstream risk.
- Assess organizational governance and vendor accountability: With risks identified and prioritized, the preparer must now determine the appropriate mitigation controls to put in place. The higher the priority of the risk, the more intense the controls. The controls fall into categories such as human oversight, review frequency, documentation standards, and deployment phasing. For the AI tool deployment and an associated risk of control dilution — where AI-accelerated volume outpaces existing review processes — controls might include an increased independent verification requirement for AI-suggested positions, as well as a phased rollout that restricts the tool to document summarization in Phase 1 and authorizes position drafting only after independent verification protocols are tested and confirmed. Given the severity of the cascade risk, the output of this step includes an organizational gap analysis of required pre-deployment controls and a vendor diligence checklist of questions that must be answered before a contract is signed.
The broader risk ecosystem
The health of any tax system depends on all the key stakeholders — in particular, taxpayers, tax authorities, and tax preparers — working in concert with one another. As the presence of AI scales up in our tax system, these stakeholders should collaborate in assessing and managing the associated risk. While tax authorities and preparers warrant distinct registers and frameworks, there are common risk areas.
The risks that appear in both registers are highlighted in the Venn diagram in Figure 5, “Shared Risks Across the Tax Ecosystem.” Tax authorities and preparers can and should work together on approaches to assess and mitigate these risks. In particular, an AI vendor may be marketing a product to both preparers and authorities that exhibits a technical deficiency, leading to an increased risk of hallucination. Raising awareness will benefit all players, including the vendor, which gains feedback to make necessary technical adjustments.
Figure 5. Shared risks across the tax ecosystem
Five risks appear in both the Tax Authorities and Tax Preparers registers, creating the potential for cross-sector harm propagation.

Innovate, but responsibly
Recommendations for authorities and preparers could be summed up briefly in these terms: With AI, you don’t have to scale fast to capture transformational results. Game–changing return on investment can result from incremental deployments — and incremental deployments give tax authorities or tax return preparers the time to get the risk diligence right. As a result, it is recommended that they always start with an assumption that an AI tool should be piloted or deployed with limited scope and tested thoroughly before it is scaled. Bounded, targeted, or segmented deployments unlock significant performance while providing time to assess how AI risks play out and the most effective methods to address them. In this way, the frameworks introduced in this article, like any effective risk management tool, are intended to enable innovation, not stop it. Stated differently, the goal is responsible innovation.
Footnotes
1 The definitions in Figure 1 and Figure 3 are intentionally concise. The complete register in Appendices A and C includes expanded definitions incorporating additional guidance.
2 Risks are currently least understood in the area of Agentic AI, and this register will need to evolve as the field develops. See, e.g., OpenAI, “Detecting and Reducing Scheming in AI Models” (Sept. 17, 2025), which investigates whether AI systems that are given increasing autonomy might generate outputs that are not merely incorrect but strategically shaped to resist correction or misrepresent their own reasoning.
3 Committee of Sponsoring Organizations of the Treadway Commission, Enterprise Risk Management — Integrating With Strategy and Performance (2017).
4 Bruckner and Coil, AI and the Modern Tax Agency: Adopting and Deploying AI to Improve Tax Administration, IBM Center for the Business of Government and American University Kogod School of Business Tax Policy Center (2024). The report was developed in consultation with IRS personnel and informed by a global roundtable of tax administrators, including a session in which the author participated as IRS commissioner.
5 The definitions in Figure 1 and Figure 3 are intentionally concise. The complete register in Appendices A and C includes expanded definitions incorporating additional guidance.
Contributor
Danny Werfel, J.D., MPP, is a former IRS commissioner. Currently, he is executive in residence at the School of Government and Policy at Johns Hopkins University in Baltimore and Polis Distinguished Fellow at the Polis Center for Politics at Duke University in Durham, N.C. After serving in 2013 as acting commissioner, from March 2023 to January 2025, he served as the 50th commissioner of the IRS. From 2006 to 2009, he served on the Federal Accounting Standards Advisory Board.For more information about this article, contact thetaxadviser@aicpa.org.
MEMBER RESOURCES
Article
Hartman, “AI Risks CPAs Should Know,” 241-2 Journal of Accountancy 52 (February 2026)
Tax Section resource
Artificial Intelligence Tax Resource Center
CPE self-study
AI at Work: Practical Tools for Professionals
For more information or to make a purchase, visit aicpa-cima.com/cpe-learning or call 888-777-7077.
