Checklist · 5 minute read
AI Release Checklist: Shipping a Change Without Surprises
An AI release may be a code change, a prompt change, a model change, or a corpus change, and all four alter behaviour. Each needs evaluation evidence, a recorded version, a cost and latency check, a staged rollout, and a rollback that someone has tested rather than documented.
Releasing an AI change is different because the change may be a prompt, a model, or a corpus rather than code. This checklist covers all four, drawn from FISTA Solutions' AI agents delivery practice.
What kinds of change need this?
Four, all of which alter behaviour.
| Change type | Why it is a release |
|---|---|
| Application code | Conventional; already covered |
| Prompt or instructions | Directly changes behaviour |
| Model or version | Changes behaviour with no code diff |
| Retrieval configuration | Changes what the model sees |
| Corpus content | Changes the answers available |
| Tool definitions | Changes what an agent does |
Evaluation evidence
The gate that distinguishes a tested change from a hopeful one. See how to build an agent evaluation harness.
- Evaluation suite run against the new configuration
- Results compared against the current production configuration
- Regressions identified and either fixed or explicitly accepted
- The case motivating the change is covered by a test
- Adversarial and edge cases included in the run
- Results attached to the change record
- The suite runs in the pipeline and blocks on regression
Version recording
An incident will ask which configuration produced an output.
- Model and version pinned explicitly, not implicit
- Prompt version recorded and deployed from version control
- Retrieval configuration versioned
- Corpus snapshot or index version recorded
- All versions logged alongside every output
- The deployed combination is reproducible from the record
- Version changes appear in the change log
Cost and latency impact
Quality improvements that triple cost are a decision, not an accident.
- Token usage measured before and after
- Cost per task computed for the new configuration
- Projected monthly cost change calculated
- Latency measured at the tail, not the average
- Time to first token checked for streaming interfaces
- Any material increase approved by the budget owner
- Provider rate limit headroom rechecked if volume changes
Staged rollout
Behaviour changes evaluation missed surface here. See feature flags guide.
- The change is behind a flag or percentage rollout
- Internal users first, then a small percentage
- A watch period defined at each stage
- Quality, cost, and latency signals watched during the period
- Rollout criteria agreed before starting
- Someone is responsible for watching, not just for deploying
- The rollout can be halted and reversed at any stage
Rollback
Tested, not documented.
- Rollback is a single action, not a rebuild
- The previous configuration remains available
- Rollback has been exercised recently
- On-call can perform it without the release engineer
- Rollback criteria agreed in advance with thresholds
- Any data written under the new configuration is identifiable
- Time to complete a rollback is known
Communication
People affected should hear before users do.
- Change described in terms of behaviour, not implementation
- Support and operations told what changes and when
- Known behaviour differences documented for reviewers
- Users told if the change is visible to them
- The change record is findable later by someone investigating
- Post-release check scheduled rather than assumed
- Someone named as the contact during the watch period
What are the most common failures?
Treating prompt and model changes as configuration rather than releases. Shipping without an evaluation run. Untested rollback. Full rollout with no staging. And version information absent from logs, which makes investigation impossible.
Who should own this?
The team owning the system owns its releases. Model and prompt changes should go through the same approval as code, which is the main change most organisations need to make.
How often should it run?
Every change. For model version changes forced by a provider, the same checklist applies on a deadline you did not choose, which is an argument for keeping the evaluation suite ready.
What evidence should it produce?
The change record with evaluation results, version details, cost and latency measurements, and rollout notes. That set answers what changed, when, and what it did.
What about corpus updates?
They change answers without any code or prompt change, and they are the category most often released outside change control.
Adding or removing documents alters what the system can say. Significant corpus changes deserve an evaluation run and a note in the change record, particularly where the content is authoritative for customer-facing answers. See AI knowledge base quality checklist.
What should you do first?
Check whether your last prompt change went through review with evaluation evidence. If it did not, bringing prompts into the release process is the highest-return change here.
How FISTA Solutions helps
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: prompt, model, and corpus changes treated as releases with evaluation evidence attached, and versions logged against outputs so incidents can be traced, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To adapt this checklist to your environment, message FISTA on WhatsApp, or read AI rollback checklist.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What counts as a release?
Any change that alters behaviour: code, prompts, model version, retrieval configuration, or the corpus. Treating only code as a release leaves three categories of change outside change control.
02What evidence should a release carry?
An evaluation run comparing the new configuration against the current one, with regressions explained, plus cost and latency measurements if anything grew.
03Why record versions?
Because an incident asks which configuration produced a given output. Without model, prompt, and corpus versions logged alongside outputs, the answer is unavailable.
04How should rollout be staged?
A small percentage or an internal group first, with a watch period on quality and cost signals before widening. Behaviour changes that evaluation missed show up here.
05What makes rollback real?
Having done it. A documented rollback path that nobody has exercised is a hypothesis, and the moment you need it is the worst time to test it.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.