Checklist · 5 minute read
AI Rollback Checklist: Reversing a Change Under Pressure
Rollback decisions happen under pressure with partial information, which is why the criteria should be agreed beforehand. Know what to revert — code, prompt, model, and corpus are separate — identify what the bad version wrote, and communicate before users discover the problem themselves.
Rollback decisions are made under pressure with incomplete information, which is why the thinking belongs beforehand. This checklist covers it, drawn from FISTA Solutions' AI agents operational work.
What is the sequence?
Six steps, in order, under pressure.
| Step | Decision or action |
|---|---|
| Confirm impact | Is this real and user-facing? |
| Decide | Roll back or fix forward |
| Identify what changed | Code, prompt, model, or corpus |
| Revert | Single action, previous known state |
| Assess written data | What did the bad version do? |
| Communicate | Users, support, stakeholders |
Decision criteria
Agreed in advance so the decision is not a debate during an incident.
- Thresholds defined for user-facing impact
- Quality degradation threshold defined and measurable
- Cost spike threshold defined
- Who can call a rollback is named, with a deputy
- Fix-forward is the exception, with criteria
- The decision does not require the person who shipped the change
- Criteria are in the runbook, not in someone's head
Identify what changed
Reverting the wrong thing wastes the outage.
- The change log consulted for everything deployed recently
- Code, prompt, model, retrieval config, and corpus checked separately
- Provider-side model changes considered, not only your own
- Version information from logs used to confirm what was running
- Timing correlated between the change and the symptom
- Multiple simultaneous changes untangled before reverting
- The suspected cause stated explicitly before acting
Perform the rollback
A single action to a known state. See AI release checklist.
- Revert to the last known good configuration, not a partial state
- Related components reverted together where they were tested together
- Feature flag disabled where that is the faster path
- Rollback completion verified by observing the signals recover
- Caches invalidated if they hold output from the bad version
- In-flight tasks handled explicitly
- Time of rollback recorded
Assess written data
Configuration reverts; records do not.
- Records created or modified during the bad period identified
- Assessment of whether those records are wrong
- Correction plan for affected records
- External actions taken — emails sent, tickets created — identified
- Affected customers identified where relevant
- Downstream systems that consumed bad output identified
- Correction executed and verified
Communication
Early and accurate beats late and complete.
- Support and operations told immediately
- Users told if they were affected
- Stakeholders told with impact and expected resolution
- Status page or equivalent updated if you have one
- A single owner for communications during the incident
- Follow-up message when resolved
- Customers whose records were corrected told, where appropriate
After the rollback
Capture what happened while it is fresh.
- Timeline recorded while details are remembered
- The failure added to the evaluation suite as a permanent case
- Why evaluation did not catch it examined
- Detection time measured and a faster signal identified
- Rollback duration measured against expectation
- Runbook updated with anything that was unclear
- A blameless review held with the people involved
What are the most common failures?
Debating the decision during the incident. Reverting code when the model changed. Forgetting records written under the bad version. Communicating after users complain. And not adding the failure to the evaluation suite afterwards.
Who should own this?
On-call performs the rollback; the system's business owner is informed and owns the communication decision. Requiring the original engineer's involvement makes the response slower than it needs to be.
How often should it run?
Exercise rollback during ordinary operations at least quarterly. Review the criteria after every incident, since each one reveals whether the thresholds were right.
What evidence should it produce?
The incident timeline, the configuration reverted to, data remediation records, and the evaluation case added afterwards. That set demonstrates the response was controlled.
What if rollback is not possible?
Then the situation is a fix-forward under pressure, which is the worst position available and the reason rollback capability matters.
Common causes are irreversible schema changes, a corpus overwritten in place, and a model version no longer offered by the provider. Each is avoidable with planning, and each should be treated as a defect in the release process. See AI disaster recovery checklist.
What should you do first?
Agree your rollback criteria and write them in the runbook. Deciding under pressure without them is how incidents get longer.
How FISTA Solutions helps
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: rollback criteria agreed in advance and exercised during ordinary operations, with data written under a bad version identified and corrected deliberately, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To adapt this checklist to your environment, message FISTA on WhatsApp, or read AI release checklist.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01When should you roll back rather than fix forward?
When the impact is user-facing and the fix is not immediately obvious. Rollback restores a known state in minutes; investigation under pressure produces worse decisions than investigation afterwards.
02What exactly gets reverted?
Whichever changed: application code, the prompt, the model version, the retrieval configuration, or the corpus. Reverting the wrong one wastes the outage.
03What about data written by the bad version?
Identify it, assess whether it is wrong, and decide on correction. Records created or modified during the bad period do not revert with the configuration.
04Why might a model rollback need prompt changes?
Because prompts tuned for the new model may perform differently on the old one. Rolling back the model alone can leave a configuration that was never tested together.
05When should you communicate?
As soon as impact is confirmed, before users report it. A brief, accurate message early costs far less than the same message after the complaints arrive.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.