Checklist · 5 minute read
AI Disaster Recovery Checklist: Planning for Provider Failure
AI systems fail in ways conventional recovery plans do not cover: a provider outage, a model withdrawn at short notice, a quality collapse with no error, or a corpus lost. Each needs a defined fallback, and the fallback that matters most is a manual path people can still use.
AI systems fail in ways conventional disaster recovery plans do not anticipate. This checklist adds the missing scenarios, drawn from FISTA Solutions' AI enablement operational work.
What scenarios need a plan?
Six, with the manual path underpinning all of them.
| Scenario | Required response |
|---|---|
| Provider outage | Failover or graceful degradation |
| Sustained rate limiting | Queue, shed load, or degrade |
| Model withdrawn | Migration under deadline |
| Quality collapse | Detection, then rollback |
| Corpus or index loss | Restore from backup |
| Total AI unavailability | Manual path, still working |
Provider outage and degradation
The most likely scenario and the one most worth automating. See what is a fallback chain.
- Behaviour on provider timeout and error is defined
- A secondary provider or model is configured and tested
- Failover is automatic where appropriate, manual where not
- Degraded mode defined: what still works without the model
- Users are told when the system is degraded
- Queueing for non-urgent work rather than failing outright
- Provider status monitoring feeds your alerting
Quality collapse
The failure that produces no error. See how to monitor AI quality in production.
- Quality metrics monitored continuously, not only at release
- Alert thresholds defined on quality, not only on errors
- Production sampling reviewed by a person on a schedule
- Rollback to a previous model or prompt version is one action
- Rollback has been tested, not only documented
- A decision threshold agreed for disabling the feature
- User-reported quality problems reach the team quickly
Model deprecation
A deadline set by someone else. See how to run a model migration.
- Deprecation notices subscribed to and monitored
- Current model versions inventoried across all systems
- Evaluation suite ready to run against a replacement
- Migration effort estimated before it is needed
- Capacity reserved for periodic migration work
- An abstraction layer makes switching a configuration change
- A worked example of a past migration documented
Data and index recovery
Several of these exist only in a vendor's system by default.
- Source corpus backed up independently of the index
- Index rebuild from source tested and timed
- Evaluation cases and results backed up
- Prompts and configuration in version control
- Embeddings recoverable or re-computable within an acceptable window
- Vendor-held data exported on a schedule
- Restore tested, not just backup
The manual path
The fallback underneath everything else.
- The manual procedure is documented and current
- Enough people still know how to perform it
- Access to the systems needed for manual work is retained
- Capacity to handle the volume manually is understood
- A threshold defined for switching to manual
- The switch has been rehearsed
- Customer communication for extended degradation is prepared
Plan maintenance
An untested plan is a hypothesis.
- Each scenario has a named owner
- Runbooks exist and are reachable during an outage
- Contact paths for the provider are known and current
- Recovery objectives defined per system and agreed with the business
- A rehearsal is scheduled at least annually
- Lessons from real incidents fold back into the plan
- The plan is reviewed when the architecture changes
What are the most common failures?
Planning for infrastructure failure and not provider failure. No quality alerting. Untested rollback. A manual path nobody can perform any more. And backups of the index but not of the source corpus.
Who should own this?
The system's business owner agrees the recovery objectives; engineering owns the technical fallbacks; operations owns the runbooks and the rehearsal. The manual path is owned by the function that used to do the work.
How often should it run?
Reviewed quarterly, rehearsed at least annually, and updated whenever the architecture or the provider changes. Rollback should be exercised more often than that — ideally during ordinary operations.
What evidence should it produce?
Rehearsal records, rollback test results, restore timings, and the current runbooks. That set demonstrates the plan is real rather than aspirational.
How much of this applies to a low-risk system?
Scale it to consequence. An internal tool that saves time needs a documented degraded mode and nothing more; a system customers depend on needs the full set.
The item that applies at every level is the manual path, because the question of what happens if this stops is asked about every system eventually. See how to staff an AI support rotation.
What should you do first?
Ask whether anyone has tested rolling back your production AI system to a previous version. If not, that is an afternoon's work and the most useful item here.
How FISTA Solutions helps
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: quality alerting alongside availability alerting, rollback exercised during ordinary operations, and a manual path kept current and rehearsed, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To adapt this checklist to your environment, message FISTA on WhatsApp, or read what is a fallback chain.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What is the most likely failure?
A provider outage or sustained rate limiting. It is common enough to plan for concretely, and the response — fall back to another provider or degrade gracefully — has to be built in advance.
02Why is quality collapse hard to detect?
Because the system keeps responding. Availability monitoring stays green while outputs become wrong, and nothing alerts unless quality itself is monitored.
03Why does the manual path matter?
Because it is the ultimate fallback. If the people who used to do the work have been reassigned and the procedure is undocumented, an extended outage becomes a business interruption rather than a degradation.
04Is model deprecation a disaster?
It is a scheduled one. A withdrawal with short notice forces migration on someone else's timetable, and teams without evaluation infrastructure cannot do it safely at speed.
05What needs backing up?
The corpus, the index, evaluation cases and results, prompts, and configuration. Several of these live only in a vendor's system unless deliberately exported.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.