Playbook ┬╖ 6 minute read
How to Run an AI Governance Review That Finds Things
A governance review finds things when it samples real systems and checks whether the evidence exists, rather than confirming that policies exist. Reviews that examine documentation find documentation, and the gap between what a policy requires and what systems actually do is where the risk lives.
Governance reviews that check policies find that policies exist. The useful review samples real systems and asks whether the evidence the policy requires actually exists. This playbook covers running one, drawing on FISTA Solutions' AI enablement work. This article is general guidance, not legal advice.
When is this worth doing?
Periodically for any organisation with AI systems in production, and before a regulatory examination where one is anticipated.
It is also worth running the first one early, while the estate is small enough that the findings are manageable and the practices can still be established rather than retrofitted.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Test the inventory | Against reality, not the register |
| 2. Sample systems | Including unnominated ones |
| 3. Request the evidence | What exists, not what is described |
| 4. Check policy against practice | The gap is the finding |
| 5. Record findings with owners | And dates |
| 6. Verify remediation | At the next review |
Step 1 тАФ Test the inventory first
Take the inventory and check it against reality: do the listed systems exist, do they have the stated owners, and are there systems in production that are not listed.
The last question is the important one. Network logs, expense records, and asking teams directly all surface systems the register does not contain, and an incomplete inventory undermines every other control.
An inventory that cannot be tested is not an inventory. See what is an ai inventory.
Step 2 тАФ Sample systems, including unnominated ones
Pick three to six systems across different teams, and include at least one that nobody offered.
Self-selected examples are the best ones. A review examining only those describes the organisation at its most prepared, which is not the question anyone is actually asking.
Weight the sample towards systems that make or influence decisions about people, which is where the consequence and the regulatory attention concentrate.
Step 3 тАФ Request the evidence rather than descriptions
For each sampled system: the inventory entry with owner, the risk assessment, evaluation results with dates and versions, approval records, oversight design, and incident history.
Ask to see them rather than asking whether they exist. The difference in answer is substantial and it is the entire value of the review.
Time it too. Evidence produced within an hour is evidence that exists as a by-product of how the system is run; evidence produced after three days was assembled for you.
Step 4 тАФ Check policy against practice
Where the policy requires something the sampled systems do not do, that is a finding тАФ and a more serious one than an absent policy.
A policy requiring impact assessments with no assessments performed documents that the organisation knew what was required and did not do it. That is worse than having no policy, and it is a common finding.
The remedy is sometimes to change the policy rather than the practice. A requirement nobody follows because it is disproportionate should be revised rather than enforced.
Step 5 тАФ Record findings with owners and dates
Each finding: what was found, why it matters, who owns the remediation, and by when.
Findings recorded as observations get noted and appear again at the next review. That pattern is how governance reviews become rituals that consume time and change nothing.
Keep the list short. Three findings that get remediated are worth more than fifteen that get filed, and the discipline of choosing produces better prioritisation.
Step 6 тАФ Verify remediation at the next review
Start the next review by checking the previous findings.
That single practice determines whether the review process has teeth. An organisation where findings persist across reviews has a governance problem larger than any individual finding, and naming it is more useful than adding to the list.
Report the remediation rate alongside the new findings. It is the most informative number the review produces.
Who should conduct it?
Someone independent of the systems being reviewed: internal audit, a different team, or an external party.
Self-review produces optimistic results reliably, not through dishonesty but because people cannot see the assumptions they share. The independence is what makes the evidence-based approach work.
The reviewer needs enough technical understanding to know what evidence to ask for. Reviews conducted entirely by non-technical assurance functions accept documents that do not demonstrate what they appear to.
How does this relate to regulatory examination?
It is a rehearsal. The questions a regulator asks тАФ what systems exist, how were they assessed, what testing was done, who approved, what happened when something went wrong тАФ are the questions this review asks.
An organisation that passes its own review comfortably will find an examination a retrieval exercise. One that cannot produce the evidence internally will not produce it for a regulator either. See how to respond to an ai regulator inquiry.
Who needs to be involved?
An independent reviewer with technical understanding, the owners of the sampled systems, and someone senior enough to assign remediation.
Without the last role, findings become recommendations and the review becomes a report.
How long does it take?
One to two weeks for a sample of a handful of systems, including the evidence requests and the write-up. Reviews taking months are usually surveying rather than sampling.
What are the common failure modes?
Checking policies. Reviewing only nominated systems. Accepting descriptions. Long finding lists. No owners or dates. And not verifying previous remediation.
How do you know it worked?
Findings that get remediated between reviews, an inventory that survives testing, evidence produced within hours rather than days, and a shrinking gap between policy and practice.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
Take your inventory and check whether any production AI system is missing from it. That single test predicts most of what the review will find.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: reviews that sample real systems and test for evidence rather than policy, findings recorded with owners and dates and verified at the following review, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read AI compliance audit cost.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why do governance reviews find nothing?
Because they examine policies rather than practice. A review confirming that an AI policy exists has confirmed a document, and the gap between what the policy requires and what systems actually do is where the risk sits.
02What should be sampled?
Real systems, including at least one nobody nominated. Self-selected examples are the best ones, and a review that only examines those describes the organisation at its most prepared.
03What evidence should exist?
For each sampled system: an inventory entry with an owner, a risk assessment tied to its use, evaluation results with dates, approval records, and incident history. Whatever cannot be produced quickly is the finding.
04What makes a finding actionable?
An owner, a date, and a specific remediation. Findings recorded as observations get noted and repeated at the next review, which is how governance reviews become rituals.
05How often should this run?
Quarterly at most for a large estate, annually for a small one. More frequent reviews end up measuring the review process rather than the systems, because remediation takes longer than a quarter and the real work happens between reviews rather than during them.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.