
Written by:
Editorial Team
DSG.AI
Agentic AI deployment in internal audit functions grew from 11% to 25% between Q1 and Q4 2025, according to Fieldguide's industry tracking data. That rate of adoption is fast for a profession that spent a decade debating whether automated workpaper tagging was a good idea.
The more important statistic is the gap between that 25% and what "deployed" actually means in most organizations. A majority of those deployments are automating single-step tasks: document classification, engagement letter generation, scheduling reminders. These are workflow improvements, not workflow transformations. The functions reaching genuine audit cycle reduction, measured as 50%+ fewer days from fieldwork start to report issuance, are doing something categorically different.
This piece is about what that difference looks like in production, and how internal audit functions get there from where most of them are today.
What most "AI in audit" deployments are
Let's be precise about what the vendors are selling and what most teams are buying.
Level 1: Task Automation. A single step in the audit workflow is handled by AI. Examples: automatically routing incoming evidence files to the right control in a workpaper; parsing engagement emails to extract dates, names, and action items; generating a first draft of an engagement letter from a template. These improve the speed and consistency of individual tasks. They do not change how much time an audit costs.
Level 2: Multi-step Workflow. A connected sequence of steps is automated. Examples: receiving a sample population from the ERP, testing each item against a defined control criterion, flagging exceptions, and drafting a finding with the exception count and supporting evidence attached. This is where meaningful time reduction begins. A manual task that takes a staff auditor 6 hours can run in under 30 minutes. The auditor's time shifts to reviewing the output, not producing it.
Level 3: Agentic Control Testing. An AI agent executes a full control test from data extraction to documented finding, including exception identification, follow-up queries for missing evidence, and workpaper closure. The audit team defines the test parameters, reviews the findings, and provides judgment on materiality and tone. The AI performs the work. This is the model operating at 250+ production deployments.
Most internal audit functions are in Level 1. The ones writing about "50% cycle time reduction" are at Level 2 or Level 3. The difference between Level 1 and Level 3 is not primarily a technology difference; the tooling for Level 3 has been available for two years. The difference is whether the audit team has designed its evidence collection and testing workflows in a form that AI systems can execute, or whether the workflows still depend on human judgment at every step because nobody ever wrote down the decision criteria.
The four workflows that generate the most time savings
Not all audit work is equally automatable. Start with the workflows where the decision criteria are explicit and the evidence sources are machine-readable. These four generate the most time savings across the broadest range of audit programs:
1. Evidence collection from source systems
Control testing typically starts with evidence requests: the back-and-forth with control owners to collect screenshots, exports, and logs. This process consumes 30-40% of field hours in a typical IT general controls engagement, per the IIA's 2024 Global Internal Audit Common Body of Knowledge. Automating it requires: (a) direct integrations to source systems (ERP, IAM, HR systems), and (b) defined logic for what constitutes sufficient evidence for each control. Organizations that have built these integrations report field hour reductions of 50-70% on ITGC testing.
2. Population testing and exception identification
The shift from sampling to full-population testing, enabled by AI processing of complete transaction sets, eliminates the statistical uncertainty inherent in sample-based audit work. Full population testing is now cost-competitive with sampling for any control tested against a machine-readable data source. The AI runs every transaction against the control criterion, identifies all exceptions (not a projected sample), and documents the count. Exception review remains with the auditor; population-level pattern identification does not.
3. Workpaper documentation
AI systems that execute test procedures can document what they did and what they found in a standardized workpaper structure. This is not a "draft a report" prompt to a general-purpose LLM. It is a structured documentation layer built on top of the test execution: test objective, population parameters, sample or full-population flag, exception count, exception details with evidence links, and a draft conclusion statement. The auditor reviews and modifies the conclusion. The mechanical workpaper production is removed.
4. Continuous monitoring with alert routing
For key financial controls with high-frequency transactions (revenue recognition, accounts payable approval, access provisioning), continuous monitoring replaces periodic testing. The AI monitors every transaction against defined control criteria and routes exceptions to the responsible auditor in real time. Monthly or quarterly testing of these controls, with sampling, becomes a maintenance artifact rather than a meaningful assurance activity.
Why most implementations stall at Level 1
The most common reason internal audit functions plateau at task automation is that the underlying workflows were never designed for machine execution. Evidence collection depends on auditors knowing which specific data points they need and which system holds each. Control testing criteria exist in the minds of experienced auditors but not in a form the AI can apply. Exception thresholds are judgment calls made individually, not documented rules.
This is the actual problem to solve before deploying AI. It is not a technology problem; it is a workflow documentation problem. The organizations that have reached Level 3 spent 3-6 months before their first AI deployment documenting: which systems hold which evidence for which controls; what constitutes a pass, a fail, and an exception requiring judgment; and which controls have enough volume and machine-readable data to test completely versus which require human evaluation of each item.
The documentation work is also the audit quality work. An audit function that can specify its testing criteria precisely enough for an AI to execute them has, in the process, closed the interpretive inconsistency that produces false negatives in manual testing.
What the audit team does instead
The objection that usually surfaces at this point: "If AI does the testing and workpaper production, what do auditors do?"
The answer, in production: auditors define the control criteria, review exceptions, exercise judgment on materiality, communicate findings, and manage the relationship with the business. They also maintain the AI systems: updating test logic when controls change, reviewing model outputs for drift, and documenting governance of the AI systems themselves. For audit teams that have done this well, the governance of the audit AI has become a recognized competency that boards and audit committees ask about.
The model is auditors governing AI systems that perform the work, not auditors performing the work themselves. The AI audit agent governance framework covers what that governance looks like in practice.
What does not happen: auditor headcount does not automatically reduce. Coverage increases instead. A function that previously tested 200 controls per year, sampling 25 per control, now tests all 200 controls, every transaction, every cycle. The audit plan expands. The function covers more risk, not the same risk with fewer people.
Getting from Level 1 to Level 2
The practical path from task automation to workflow automation is not buying a different tool. It is building the evidence integration layer and the control-criterion documentation that allow AI systems to execute tests rather than assist humans performing them.
Three first steps that work in practice:
Audit one high-frequency IT control from data extraction to documented finding. Pick a control tested against a machine-readable data source: user access review, password policy compliance, change management ticketing. Build the evidence integration directly from the source system, document the test criteria precisely, and execute the test with full-population coverage. Run it alongside your existing manual process for one cycle to validate. Then retire the manual test.
Document decision criteria for your ten most common exceptions. For each exception type, write the criteria that would cause an AI to route an exception to a senior auditor versus close it as clearly immaterial. This is the judgment codification work that makes Level 3 possible.
Map which of your annual engagements have machine-readable evidence for at least 70% of tested controls. Those are your Level 2 targets for the next planning cycle. Engagements with primarily interview-based or qualitative evidence are not AI candidates yet.
At DSG, assureIQ (deployed across 40+ enterprise clients) runs all three of these steps as the onboarding process. The result is an audit function that tests controls at full population, closes engagements in half the previous cycle time, and covers 3-5x the original risk universe. The 50%+ cycle time reduction and 3-5x coverage increase are measured outcomes, not projections.
The technology is not the constraint. The workflow design is.
<!-- related-links:start (auto-managed by seo/sync-internal-links.mjs) -->Related
- Internal Audit Sourcing Cost Reference (2026): In-House vs. Mid-Tier vs. Big 4 vs. AaaS
- AI in Internal Audit: Adoption Statistics and Research (2026 Library)
- Internal Audit Engagement Benchmarks: Cycle Times, Hours, and Rates by Provider Tier (2026)
- Audit-as-a-Service: What It Is, What It Costs, and When It Beats Hiring
- Co-Sourcing vs. Outsourcing Internal Audit: Decision Framework and Real Costs
- What Compliance-as-a-Service Actually Includes (and What Vendors Leave Out)


