FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Playbook ┬╖ 6 minute read

How to Set Up AI Change Control That Does Not Block Work

AI change control works when changes are classified by risk, only the consequential ones are gated, evidence is produced automatically by the pipeline, and routine changes stay fast. Processes that treat every prompt edit as a formal change get bypassed within a month.

By FISTA Solutions┬╖ AI-Native Engineering Team┬╖
How to Set Up AI Change Control That Does Not Block Work article cover

Change control for AI has to cover prompts, models, and corpora that conventional processes ignore, while staying fast enough that people use it. This playbook covers both halves, drawing on FISTA Solutions' AI enablement work.

When is this worth doing?

When an AI system is in production and being changed regularly, and particularly where validation, model risk, or regulatory requirements apply.

It is not worth imposing on a prototype. Change control on something still being built slows discovery without protecting anything.

What does the sequence look like?

StepPurpose
1. Define what counts as a changeBroader than code
2. Classify by riskConsequence, not size
3. Gate the consequential onesLeave routine changes fast
4. Automate the evidenceFrom the pipeline
5. Define the emergency pathBefore it is needed
6. Review the classificationAs the system matures

Step 1 тАФ Define what counts as a change

List everything that alters behaviour: prompts, system instructions, model versions, retrieval corpora, chunking and embedding configuration, thresholds, tool definitions, and code.

Conventional change control covers the last one. That leaves most of the behaviour surface outside the process, which is how a system's behaviour changes with no record and nobody can say what happened.

Make the list specific to your system and publish it. Ambiguity about what needs a record produces inconsistent practice.

Step 2 тАФ Classify by risk, not by size

A one-word prompt change on a system making credit decisions is more consequential than a large refactor of an internal summarisation tool.

Classify by what the system does and who it affects: customer-facing, regulated, irreversible actions, and safety-relevant systems warrant gating. Internal, reversible, low-consequence systems do not.

That classification should be attached to the system rather than decided per change, which removes an argument from every release.

Step 3 тАФ Gate the consequential changes only

For high-risk systems: evaluation must pass, a reviewer must approve, and the change must be recorded with its versions.

For low-risk systems: automated checks and a version record are sufficient.

That split is what keeps the process usable. A process requiring a review board for every prompt tweak on an internal tool will be bypassed, and the bypassing will extend to the changes that mattered.

Step 4 тАФ Automate the evidence

Evaluation results, version records, test outcomes, and approvals should be captured by the pipeline as the change moves through.

Evidence produced by a form somebody fills in afterwards is worse in every respect: it takes longer, it is less accurate, and it describes what someone remembers rather than what happened.

This is also what makes audit cheap. A change record assembled automatically at the time is retrieval; one reconstructed later is a project. See AI compliance audit cost.

Step 5 тАФ Define the emergency path

A documented route for urgent changes: who can authorise, what minimum evidence is required, and mandatory retrospective review within a defined period.

Processes without an emergency path produce an exception culture. People invent a route under pressure, it is not recorded, and it becomes the normal route for anything inconvenient.

Track how often the emergency path is used. Frequent use means the ordinary process is too slow, which is a finding about the process rather than about the people.

Step 6 тАФ Review the classification periodically

Systems change in consequence as they become more embedded. A tool that started as an internal convenience and now informs customer decisions has moved risk category.

Review classifications on a schedule and when a system's use expands. Systems whose classification never changes are usually mis-classified in one direction or the other.

Record the reasoning so the next review starts from a known position.

How does this handle provider model updates?

As changes, which they are. A provider updating a model alters your system's behaviour, and a process that only covers your own deployments misses it entirely.

Pinning versions is what makes this controllable: the update becomes something you adopt deliberately, with evaluation, rather than something that happens to you. Systems calling unpinned aliases have an uncontrolled change process regardless of what the policy says. See how to version prompts and models.

What about changes made by non-engineers?

Content changes to a retrieval corpus are made by subject matter people, and they are behaviour changes.

The workable pattern is ownership and review dates rather than a technical change process: documents have owners, changes are attributed, and the corpus version is recorded. Requiring a change ticket for every policy update makes the corpus stale, which is worse.

Who needs to be involved?

An engineer to automate the evidence, whoever owns the system, and the governance or validation stakeholder where one exists.

Processes designed by governance alone are thorough and bypassed. Processes designed by engineering alone miss the evidence requirements.

How long does it take?

Two to four weeks to define, classify, and automate the evidence capture. The classification conversation is quick; automating the pipeline evidence takes most of it.

What are the common failure modes?

Covering only code. Gating everything. Evidence collected by form. No emergency path. Unpinned model versions. And classifications that never get reviewed.

How do you know it worked?

Changes traceable to versions and approvals, regressions attributable, emergency path used rarely and recorded, and nobody routing around the process.

What does it cost?

Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.

Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.

What should you do first?

List every way your system's behaviour can change and mark which ones currently produce a record. The gaps are the scope.

How do you measure whether the process is used?

By the proportion of behaviour changes that have a record. Compare the changes visible in your version control and pipeline against the changes visible in the system's behaviour over the same period тАФ a discrepancy means changes are happening outside the process.

The other signal is how often people ask whether something counts as a change. Frequent questions mean the definition is unclear; no questions at all usually means nobody is reading it.

How FISTA Solutions helps

FISTA Solutions runs this work alongside client teams rather than around them: change classified by consequence so routine work stays fast, evidence produced automatically by the pipeline rather than by a form, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.

To run this with support, message FISTA on WhatsApp, or read how to version prompts and models.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01What counts as a change in an AI system?

Prompt edits, model version updates, retrieval corpus changes, configuration and threshold adjustments, and tool definition changes. Conventional processes cover code and miss most of those, which is where untracked behaviour change comes from.

02Should every change be gated?

No. Gating everything makes the process slow enough that people work around it. Classify by risk and gate the consequential changes, while routine low-risk changes proceed on automated checks.

03How should evidence be produced?

By the pipeline. Evaluation results, version records, and approvals captured automatically as the change moves through produce better evidence than a form somebody fills in afterwards.

04What about emergency changes?

Define a path for them with retrospective review rather than leaving people to invent one. Processes without an emergency route produce an exception culture where the route is invented each time and never recorded.

05How do you know the process is working?

Changes have traceable records, regressions are attributable, and people are not routing around the process. If the third is happening, the first two are illusory.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.

Start a project