FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Trends · 5 minute read

The Coming Audit of AI Systems and How to Be Ready

Auditors will ask what your AI systems decided, on what basis, under whose authority, and how you knew they were working. Those questions can only be answered from records captured at the time, and most deployed systems do not capture them. Readiness is a logging decision made early.

By FISTA Solutions· AI-Native Engineering Team·
The Coming Audit of AI Systems and How to Be Ready article cover

AI systems are entering the scope of ordinary audit, and the questions asked are mundane. The difficulty is that most systems cannot answer them. This piece covers readiness, drawing on FISTA Solutions' AI enablement governance work.

What is asked and what is needed?

Six questions and the record each requires.

QuestionRecord required
What AI systems exist?Maintained inventory
What does each decide?Purpose and scope documentation
On what data?Data lineage and sources
Who approved it?Approval record with date
How do you know it works?Dated evaluation results
How was this decision made?Full interaction record

Why is the inventory the first finding?

Because scope precedes everything.

An auditor cannot assess control over systems the organisation cannot enumerate. In practice, most organisations have AI in more places than the central list shows — embedded in vendor products, built by individual teams, running in spreadsheets.

Building and maintaining the inventory is unglamorous and it is the prerequisite. Include vendor products with AI features, because those are in scope too. See AI service catalog template.

What does a decision record contain?

Enough to reconstruct the specific decision, not a summary of the process.

The input received, the context retrieved and supplied, the model and version, the output produced, the action taken, and any human review with its outcome. Linked so the whole interaction can be assembled from an identifier.

This cannot be produced later. If the record was not captured at the time, the answer is that you do not know, which is the finding. See the compliance layer of AI.

What counts as evaluation evidence?

Dated results from a systematic process, with evidence that regressions were acted on.

An evaluation suite run on every change, with results stored and a record of what happened when quality dropped, demonstrates control. An engineer saying the outputs look good does not.

This is where teams with evaluation infrastructure are dramatically better placed. The audit requirement and the engineering requirement point at the same artefact. See AI eval report template.

How is oversight demonstrated?

Through records of reviews that happened, not policies saying they should.

If humans approve certain decisions, the record should show which were approved, by whom, when, and how many were changed. An approval rate of a hundred percent with no changes suggests the review is nominal, which is itself a finding.

Designing oversight so it produces evidence — and so reviewers have the time and information to do it properly — is the substance. See human in the loop AI explained.

What about vendor systems?

They are in scope, and the records are frequently unavailable.

When a vendor's AI feature makes decisions affecting your customers, you are accountable for those decisions. If the vendor cannot supply an audit trail, you cannot answer the question.

That makes audit trail export a procurement requirement rather than a technical detail. Ask before signing, because after signing there is no leverage. See AI third party risk checklist.

What should be done now?

Capture the records, because they cannot be created retrospectively.

Start the inventory. Turn on decision logging with a defined schema. Store evaluation results with dates. Record approvals. None of it is difficult and all of it is impossible to backfill.

A year of records is the difference between answering in days and reconstructing from partial evidence over months.

What is the counter-argument?

The counter is that formal AI audit is not yet widespread and building for it is premature. Two responses: the records required are the same ones needed to operate the system well, and organisations in regulated sectors are already being asked. The cost of capturing is low; the cost of not having captured is unbounded.

What does this change for engineering teams?

It means logging and inventory are product requirements from the first deployment, not governance overhead added later.

It also means schema stability matters: records captured under one schema and queried years later must still be interpretable, which argues for a documented, versioned format.

What does this change for buyers?

It means requiring audit trail access and export in contracts, and asking how a vendor would help you answer a regulator's question about a specific decision.

A vendor without an answer has transferred a liability to you.

What should leaders do about it now?

Start the inventory this quarter and require decision logging on every deployed system. Those two actions cover most of what will be asked.

Then assign an owner for audit readiness, because it is nobody's job by default and therefore does not happen.

Does this change with agents?

It intensifies. An agent takes actions, so the audit question becomes what it did rather than what it suggested, and the trajectory matters as much as the outcome.

Logging every step, the authority under which it acted, and the limits in force at the time is the requirement. Systems logging only final actions cannot explain how they were reached. See the shift from chatbots to agents.

How will you know if this is happening?

Watch for regulators publishing AI-specific expectations in your sector, for customers asking about automated decisions, and for internal audit adding AI to its plan. Each means the question is arriving.

How FISTA Solutions reads this

FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: decision records captured at the time with a documented schema, and evaluation results stored with dates so control can be demonstrated rather than asserted, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.

To discuss what this means for your roadmap, message FISTA on WhatsApp, or read AI governance framework.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01What will auditors actually ask?

What AI systems you run, what each decides, what data it uses, who approved it, how you monitor quality, and how a specific decision was reached. All ordinary questions with unusual difficulty.

02Why is a model inventory first?

Because scope has to be established before anything else. An organisation that cannot list its AI systems cannot demonstrate control over them, and that finding precedes every other one.

03What makes decision records hard?

They must reconstruct a specific past decision: the input, the context supplied, the model version, the output, and any human involvement. A summary is not sufficient and cannot be produced retrospectively.

04What evaluation evidence is expected?

That you test quality systematically, that results are recorded, and that regressions are acted on. Assertion that the system works is not evidence; a dated evaluation report is.

05How do you demonstrate human oversight?

With records of reviews that actually occurred — what was reviewed, by whom, what they changed. A policy stating that humans review is not evidence that they did.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project