FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Leadership · 4 minute read

The Internal Audit Leader's Guide to AI Agents

Internal audit leaders should build an audit program for agents covering inventory, permissions, evaluation evidence, human oversight, monitoring, change control, and incidents; demand evidence rather than assurances; and use agents in audit work for testing and documentation while preserving the third line's independence and judgment.

By FISTA Solutions· AI-Native Engineering Team·
The Internal Audit Leader's Guide to AI Agents article cover

Internal audit encounters AI agents from two directions: as a control environment to assess, and as a set of tools that could transform audit work itself. Both matter, and the second must not compromise the first. This guide gives audit leaders an audit program for agents, the evidence to demand, and a way to use agents in audit work while keeping the third line independent.

How should the agent estate be audited?

Like any control environment, in three phases.

Phase one: inventory and completeness. Does a complete inventory of agents exist, including those embedded in vendor products? How was completeness established? Who owns each, and what risk tier is assigned? An incomplete inventory is the most common and most consequential finding, because nothing else can be assured.

Phase two: design. For a sample across tiers:

ControlWhat to test
PermissionsAre they least privilege for the agent's stated job? Who approved them?
Approval thresholdsWho set them, on what authority, and are they enforced?
SegregationCan one agent both prepare and approve? Create a vendor and pay it?
EvaluationIs there a test set of real cases with a defined pass threshold?
Human oversightWhich actions require review, and is review actually performed?
MonitoringWhat alerts exist, and who responds?
Change controlAre prompts, tools, and model versions controlled, including provider updates?
Incident proceduresIs there a tested kill switch and a response process?

Phase three: operating effectiveness. Sample traces and verify that what the design says happens actually happened; review access review records, evaluation runs over the period, change records, alert responses, and incidents with remediation.

The executive guide to AI agent governance describes the structure audit is testing.

What evidence should be demanded?

Evidence, not assertions. "The agent was tested" is not evidence; an evaluation report with a date, a case count, a pass rate, and a threshold is. "There is human oversight" is not evidence; review records showing what was reviewed and what was changed are. Audit's most useful contribution in the first year of an agent program is often simply insisting on this distinction, which raises the standard across the organization. The AI evaluation explained for executives piece explains what good evaluation evidence looks like.

What findings recur?

Incomplete inventory, especially vendor-embedded AI; permissions broader than the job requires, usually inherited from a service account; evaluation performed once before launch and never repeated, so drift is undetected; thresholds set by implementers rather than by the business owner; audit trails that record the action but not the basis for it; and kill switches that exist in documentation but have never been tested. Each is straightforward to remediate if found early.

How can audit use agents in its own work?

Substantially, and the gains are real:

  • Population testing instead of sampling: an agent can check every transaction against a control rule rather than forty.
  • Document review for policy compliance, contract terms, and evidence completeness.
  • Evidence gathering: requesting, chasing, and organizing what auditees provide.
  • Workpaper assembly and cross-referencing.
  • Continuous monitoring between audits, flagging exceptions for follow-up.

Judgment, conclusions, and the opinion remain with auditors, and agent-assisted procedures must themselves be documented and reviewable, including what the agent did and how its reliability was assessed.

How is independence preserved?

By separating building from auditing. Internal audit should not build or own the agents it audits; tools used in audit work should be provided and maintained by IT or the second line under audit's specification, or procured, with audit documenting its own procedures. Where audit uses an agent to test a control, someone other than the tool's user should assess its reliability. This is the same logic applied to any audit tool, and it matters more here because the tools are new and their failure modes are unfamiliar.

What should audit tell the committee?

Coverage of the agent estate; findings by theme and severity; management's remediation status; the maturity trajectory of the control environment; and audit's own use of agents with the independence safeguards applied. The AI oversight and fiduciary duty piece covers what the committee is accountable for.

What should internal audit leaders ask?

  • Is the inventory complete, and how do we know?
  • For a sampled agent, can management produce the evaluation evidence and the trace?
  • Have permissions been reviewed since deployment?
  • When was the kill switch last tested?
  • Which agents changed model versions this year, and was that a controlled change?

How can FISTA Solutions help internal audit?

FISTA Solutions builds AI agents whose permissions, evaluation evidence, traces, access reviews, and change records are designed for assurance from the start, and works with audit and risk functions through its AI enablement practice to define the evidence standard before deployments rather than after. Since 2017, FISTA has delivered 150+ projects for 50+ companies across 12+ countries.

To review the auditability of agents already running in your organization, talk to FISTA on WhatsApp, or read AI audit and accountability.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01How should internal audit audit AI agents?

Start with the inventory and completeness of it, then test design: permissions, approval thresholds, evaluation before release, human oversight, monitoring, change control, and incident procedures. Then test operating effectiveness by sampling traces, access reviews, evaluation runs, and change records over the period.

02What evidence should internal audit demand about AI agents?

Evaluation results with dates and pass rates, complete traces for sampled transactions, permission configurations and access review records, change control records including provider model updates, monitoring alerts and responses, incident reports with remediation, and the inventory's completeness evidence.

03What are the most common AI audit findings?

An incomplete inventory, especially agents embedded in vendor products; permissions broader than the agent's job requires; evaluation performed once before launch and never repeated; approval thresholds set by implementers rather than the business; audit trails that do not capture the decision basis; and no tested kill switch.

04Can internal audit use AI agents in audit work?

Yes, for population testing rather than sampling, document review, control evidence gathering, workpaper assembly, and continuous monitoring. Judgment, conclusions, and the opinion remain with auditors, and the work performed by agents must itself be documented and reviewable.

05How does internal audit stay independent while using AI?

By not building or owning the agents it audits, using tools provided and maintained by the second line or IT under audit's specification, documenting its own agent-assisted procedures, and having someone other than the user of a tool assess its reliability. The same independence logic that applies to any audit tool applies here.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project