FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Checklist · 5 minute read

AI Release Checklist: Shipping a Change Without Surprises

An AI release may be a code change, a prompt change, a model change, or a corpus change, and all four alter behaviour. Each needs evaluation evidence, a recorded version, a cost and latency check, a staged rollout, and a rollback that someone has tested rather than documented.

By FISTA Solutions· AI-Native Engineering Team·
AI Release Checklist: Shipping a Change Without Surprises article cover

Releasing an AI change is different because the change may be a prompt, a model, or a corpus rather than code. This checklist covers all four, drawn from FISTA Solutions' AI agents delivery practice.

What kinds of change need this?

Four, all of which alter behaviour.

Change typeWhy it is a release
Application codeConventional; already covered
Prompt or instructionsDirectly changes behaviour
Model or versionChanges behaviour with no code diff
Retrieval configurationChanges what the model sees
Corpus contentChanges the answers available
Tool definitionsChanges what an agent does

Evaluation evidence

The gate that distinguishes a tested change from a hopeful one. See how to build an agent evaluation harness.

  • Evaluation suite run against the new configuration
  • Results compared against the current production configuration
  • Regressions identified and either fixed or explicitly accepted
  • The case motivating the change is covered by a test
  • Adversarial and edge cases included in the run
  • Results attached to the change record
  • The suite runs in the pipeline and blocks on regression

Version recording

An incident will ask which configuration produced an output.

  • Model and version pinned explicitly, not implicit
  • Prompt version recorded and deployed from version control
  • Retrieval configuration versioned
  • Corpus snapshot or index version recorded
  • All versions logged alongside every output
  • The deployed combination is reproducible from the record
  • Version changes appear in the change log

Cost and latency impact

Quality improvements that triple cost are a decision, not an accident.

  • Token usage measured before and after
  • Cost per task computed for the new configuration
  • Projected monthly cost change calculated
  • Latency measured at the tail, not the average
  • Time to first token checked for streaming interfaces
  • Any material increase approved by the budget owner
  • Provider rate limit headroom rechecked if volume changes

Staged rollout

Behaviour changes evaluation missed surface here. See feature flags guide.

  • The change is behind a flag or percentage rollout
  • Internal users first, then a small percentage
  • A watch period defined at each stage
  • Quality, cost, and latency signals watched during the period
  • Rollout criteria agreed before starting
  • Someone is responsible for watching, not just for deploying
  • The rollout can be halted and reversed at any stage

Rollback

Tested, not documented.

  • Rollback is a single action, not a rebuild
  • The previous configuration remains available
  • Rollback has been exercised recently
  • On-call can perform it without the release engineer
  • Rollback criteria agreed in advance with thresholds
  • Any data written under the new configuration is identifiable
  • Time to complete a rollback is known

Communication

People affected should hear before users do.

  • Change described in terms of behaviour, not implementation
  • Support and operations told what changes and when
  • Known behaviour differences documented for reviewers
  • Users told if the change is visible to them
  • The change record is findable later by someone investigating
  • Post-release check scheduled rather than assumed
  • Someone named as the contact during the watch period

What are the most common failures?

Treating prompt and model changes as configuration rather than releases. Shipping without an evaluation run. Untested rollback. Full rollout with no staging. And version information absent from logs, which makes investigation impossible.

Who should own this?

The team owning the system owns its releases. Model and prompt changes should go through the same approval as code, which is the main change most organisations need to make.

How often should it run?

Every change. For model version changes forced by a provider, the same checklist applies on a deadline you did not choose, which is an argument for keeping the evaluation suite ready.

What evidence should it produce?

The change record with evaluation results, version details, cost and latency measurements, and rollout notes. That set answers what changed, when, and what it did.

What about corpus updates?

They change answers without any code or prompt change, and they are the category most often released outside change control.

Adding or removing documents alters what the system can say. Significant corpus changes deserve an evaluation run and a note in the change record, particularly where the content is authoritative for customer-facing answers. See AI knowledge base quality checklist.

What should you do first?

Check whether your last prompt change went through review with evaluation evidence. If it did not, bringing prompts into the release process is the highest-return change here.

How FISTA Solutions helps

FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: prompt, model, and corpus changes treated as releases with evaluation evidence attached, and versions logged against outputs so incidents can be traced, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.

To adapt this checklist to your environment, message FISTA on WhatsApp, or read AI rollback checklist.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01What counts as a release?

Any change that alters behaviour: code, prompts, model version, retrieval configuration, or the corpus. Treating only code as a release leaves three categories of change outside change control.

02What evidence should a release carry?

An evaluation run comparing the new configuration against the current one, with regressions explained, plus cost and latency measurements if anything grew.

03Why record versions?

Because an incident asks which configuration produced a given output. Without model, prompt, and corpus versions logged alongside outputs, the answer is unavailable.

04How should rollout be staged?

A small percentage or an internal group first, with a watch period on quality and cost signals before widening. Behaviour changes that evaluation missed show up here.

05What makes rollback real?

Having done it. A documented rollback path that nobody has exercised is a hypothesis, and the moment you need it is the worst time to test it.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project