How to Govern an AI Audit Agent: The Explainability Gap No One Is Solving

Written by:

E

Editorial Team

DSG.AI

EY deployed enterprise-scale AI audit agents in early 2026. Diligent unveiled AuditAI at the IIA GAM conference in March 2026, then followed with an "Agentic GRC Workforce" announcement at its Elevate conference in April. The category has landed. What no vendor is discussing clearly is how internal audit functions should govern the AI systems they are about to deploy to perform their work.

This is not the question of how to govern AI risk in the organization (a CISO question). It is the question of how a CAE governs the AI systems running the internal audit function itself, where the outputs are audit findings, control assessments, and assurance opinions delivered to the board. The governance requirements are different, and most existing AI governance frameworks were not written with audit independence in mind.

Why AI Audit Agent Governance Is Different

Most enterprise AI governance frameworks address how the organization manages AI systems it deploys to customers or uses in operational processes. The key concerns are fairness, accuracy, safety, and regulatory compliance.

AI audit agents introduce a different set of concerns. The audit function is an independent assurance provider. Its outputs are formal opinions about the organization's control environment. If an AI system performs audit procedures, generates findings, or classifies control failures, the audit function's independence and professional judgment requirements extend to the AI performing that work.

The IIA's International Standards for the Professional Practice of Internal Auditing (the Standards) require that internal audit findings be based on sufficient, reliable, relevant, and useful evidence. The 2024 Standards update introduced explicit language on AI tool use. The CAE bears accountability for whether evidence gathered by an AI system meets those criteria, regardless of the vendor's marketing claims.

The explainability gap is specific: most AI audit agent vendors describe their systems in terms of what the agent does (collects evidence, tests controls, classifies findings). Fewer describe how the system reached a specific finding conclusion in a specific engagement, and fewer still provide the CAE with a documented audit trail that would satisfy an external review.

The Three Risks Most CAEs Are Not Accounting For

Risk 1: The agent produces findings the auditor cannot explain.

AI systems that classify control test results into "pass," "fail," or "exception" categories are making what is legally and professionally a judgment call. If an agent classifies a segregation of duties exception as "low risk" based on a pattern-matching model and that classification is wrong, the CAE is accountable for the conclusion, but the agent has no documentation of how it reached it.

Ask your vendor: for a specific finding generated by the AI agent in a past engagement, can you produce a document that explains why the system reached that conclusion and what evidence it relied on? If the answer is "the model produces a confidence score but not an explanation," you have an explainability problem that will matter in the next external quality assessment.

Risk 2: Model drift changes what "pass" means over time.

AI agents trained on control testing data learn patterns from prior audit cycles. As the organization's control environment changes (new systems, new controls, staff turnover, process changes), an agent trained on prior-period data may apply outdated pass/fail thresholds without alerting the audit team that its reference frame has shifted.

Auditor judgment is calibrated to the current environment. An AI model is calibrated to its training period. Continuous retraining is a technical requirement, not a feature.

Risk 3: The agent's decisions are outside the audit committee's visibility.

Audit committees receive findings and conclusions. If those conclusions were generated by an AI agent rather than a trained auditor, the audit committee has a legitimate question about the independence and professional judgment applied. Most CAEs deploying AI audit agents are not proactively disclosing this in their audit committee communications.

The IIA's guidance on the Chief Audit Executive role requires transparency with the board about significant changes in the audit approach. Shifting 30-40% of control testing to AI agents is a significant change in approach. The absence of disclosure is a governance gap.

Twelve Questions to Ask Before Deploying an AI Audit Agent

#QuestionWhat it reveals
1Can the agent produce a finding-level explanation (not just a confidence score) for each audit conclusion?Explainability for external review
2Is the agent's decision logic documented in a technical specification the audit function owns?Independence of the audit trail
3How often is the model retrained, and who decides when retraining is triggered?Drift management
4What is the agent's false negative rate on control testing (failing to detect a real failure)?Audit risk quantification
5How does the agent handle novel control environments it was not trained on?Boundary behavior
6Does the agent's output satisfy the IIA's criteria for sufficient, reliable evidence?Standards compliance
7Is the agent's activity logged in a way that supports an external quality assessment?QA readiness
8What human review step occurs before an AI-generated finding enters the formal audit report?Oversight architecture
9How are the agent's errors tracked and fed back into model improvement?Learning mechanism
10Who has access to the agent's audit trail, and is it stored separately from the vendor's system?Data sovereignty
11Does the vendor's contract address what happens to the agent's trained model if the contract ends?IP and continuity risk
12Has the AI system itself been audited against an AI management framework (ISO 42001 or equivalent)?AI-of-AI governance

Question 12 is the one most vendors will avoid. The ISO 42001 standard for AI management systems was designed exactly for this situation: an organization deploying an AI system in a high-stakes assurance context needs documentation of how that AI system is managed, monitored, and controlled. A vendor whose AI audit agent is not documented to ISO 42001 or equivalent is asking the CAE to run audit procedures through a system that has not been audited.

What Good Governance Looks Like in Practice

A well-governed AI audit agent deployment includes:

Before deployment: A documented decision about which audit procedures the agent will and will not perform. The agent tests controls; a human auditor reviews exceptions and forms the audit opinion. Clear delineation of what "the agent does" versus "the auditor concludes."

During operation: A finding review protocol where every AI-generated exception is reviewed by a credentialed auditor before it is included in a final report. The review should be documented: auditor name, review date, and any modification to the agent's initial classification.

At audit committee reporting: Disclosure that AI-augmented procedures were used in the audit cycle, the scope of those procedures, and the human oversight applied. Most audit committees will welcome this transparency; the ones that don't are telling you something about their tolerance for novelty in the assurance function.

Annually: An independent assessment of the agent's performance against the prior year's false negative rate benchmark, with results reported to the audit committee alongside the internal audit quality metrics.

The AI agents for internal audit piece covers what agents currently do well and where human judgment remains necessary. This piece is about the governance layer on top: who is accountable when the agent is wrong, and how you know.

The CAE's Accountability Does Not Transfer to the Vendor

The final point is the one that matters most. Vendor contracts for AI audit agent software do not and cannot transfer professional accountability for audit conclusions to the vendor. The CAE is the responsible party under the IIA Standards and under the organization's governance framework.

Deploying an AI audit agent shifts the operational burden of evidence collection and control testing to the system. It does not shift the professional judgment burden. The CAE who deploys an AI agent that generates a false negative on a material control failure is accountable for that finding under the same standards that apply to any audit quality failure.

The practical implication: govern the agent like you govern a junior auditor. Review its work. Document your review. Correct its errors. Report its performance. That governance framework is not glamorous, but it is what the Standards require, and it is what an external quality assessment will look for.

Diligent's "Agentic GRC Workforce" framing positions AI agents as headcount replacement. That framing is not a governance framework. A CAE who governs an AI agent like headcount will pass external review. One who does not will not.


Sources:

<!-- related-links:start (auto-managed by seo/sync-internal-links.mjs) -->

Related

<!-- related-links:end -->