Playbook ¡ 6 minute read
How to Build an AI Runbook Someone Can Use at 3am
An AI runbook is useful when it covers the incidents these systems actually produce â quality degradation, provider failures, unsafe output, cost spikes â with each entry short enough to follow under pressure and tested by someone who did not write it.
Runbooks written from imagined failures do not help during real ones. This playbook covers the incidents AI systems actually produce and how to write entries somebody can follow at three in the morning, drawing on FISTA Solutions' AI agents work.
When is this worth doing?
Before an AI system reaches production, and updated after every incident thereafter.
A system going live without a runbook is a system whose first incident will be handled by improvisation, and the improvisation will be done by whoever is awake rather than whoever knows the system.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. List the incident types | The ones that actually happen |
| 2. Write each entry short | Detect, check, act, escalate |
| 3. Link from the alerts | Findable under pressure |
| 4. Test with an outsider | Someone who did not write it |
| 5. Update after incidents | What was useful, what was missing |
| 6. Prune what is unused | Short runbooks get read |
Step 1 â List the incidents that actually happen
Quality degradation, provider outage, rate limiting, unsafe or inappropriate output, cost spike, escalation overload, retrieval returning nothing useful, and conventional availability failures.
The first seven are AI-specific and absent from most runbooks. Quality degradation in particular is the characteristic failure of these systems and the one nobody has a procedure for.
Start from your own incident history if you have one. Imagined failure lists are systematically different from real ones.
Step 2 â Write each entry short and structured
How it is detected, what to check first, immediate mitigation, who to involve, what to record, and when to escalate.
Four to six steps. Entries written as narrative get skimmed under pressure, and skimming a narrative produces a partial understanding rather than an action.
Put the mitigation before the diagnosis. Someone at three in the morning needs to stop the bleeding first and understand it afterwards, and a runbook that leads with investigation gets the order wrong.
Step 3 â Link the runbook from the alert
Every alert should carry a link to its runbook entry.
Runbooks that require finding the right page in a wiki get used less than ones that arrive with the alert. That sounds trivial and it is the difference between a runbook that is consulted and one that exists.
If an alert has no runbook entry, either write one or question whether the alert should fire. Alerts nobody knows how to respond to are noise by definition.
Step 4 â Test with someone who did not write it
Run a drill: pick an incident type, give it to someone unfamiliar, and watch them follow the entry.
This finds the assumed knowledge â the dashboard nobody named, the command whose syntax is wrong, the person whose contact details are stale. Authors cannot find those because they know the answers.
Do it for the high-severity entries at least. Those are the ones where following the runbook correctly matters most and where improvisation is most expensive.
Step 5 â Update after every incident
Add what was actually useful, remove what was not, and add the incident type if it was not covered.
Runbooks age quickly. Systems change, alerts change, and people leave, and an entry that referenced a dashboard that no longer exists is worse than no entry because it wastes time.
Make the update part of the postmortem rather than a separate task. Tasks that follow an incident and are not part of a process do not get done. See how to run an ai incident postmortem.
Step 6 â Prune aggressively
Remove entries for incidents that have never occurred and that nobody can imagine occurring.
Long runbooks get read less. A runbook with eight entries covering what actually happens is more useful than one with forty covering every theoretical failure, because the eight are findable.
The same applies within entries: remove steps nobody used during real incidents.
What belongs in the quality degradation entry?
Check whether the model version changed, whether the retrieval corpus changed, whether a deployment correlates, whether the input distribution shifted, and what the evaluation suite says.
Then: how to roll back a prompt or model change, how to disable a feature, and how to route to a fallback. Those are the mitigations available, and having them listed removes the hesitation about whether the responder is allowed to use them.
This is the entry most organisations do not have and most need.
Who should be able to follow it?
Anyone on the rotation, including people who did not build the system.
That is the design constraint. Runbooks written for the team that built the system encode assumed context, and the system's first out-of-hours incident is frequently handled by someone else.
Name the escalation contact explicitly, with a second name. Single points of contact are unavailable exactly when they are needed.
Who needs to be involved?
Whoever operates the system, whoever built it, and someone from the rotation who will test it.
The tester is the role that makes the runbook real rather than aspirational.
How long does it take?
A day to write a first version covering the main incident types, then continuous updating. Runbooks that take weeks to produce are usually too long to use.
What are the common failure modes?
Covering only availability. Narrative entries. Not linked from alerts. Never tested. Never updated. And accumulating entries for failures that do not occur.
How do you know it worked?
An unfamiliar responder handling an incident from the runbook, entries updated after incidents, and the runbook staying short enough that people read it.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
Write the quality degradation entry. It is the incident these systems produce most often and the one least likely to already be covered.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: runbooks covering the failures AI systems actually produce, entries tested by someone who did not write them before they are needed, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read how to set up AI on-call.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What incidents should a runbook cover?
Quality degradation, provider outage and rate limiting, unsafe or inappropriate output, cost spikes, escalation overload, and conventional availability failures. Conventional runbooks cover the last one and none of the others.
02What should each entry contain?
How the incident is detected, what to check first, immediate mitigation, who to involve, what to record, and when to escalate. Four to six steps, not a narrative.
03How long should an entry be?
Short enough to follow under pressure at an unfamiliar hour. A page at most. Entries that require reading two screens before the first action get skipped in favour of improvisation.
04How do you know it works?
Someone who did not write it follows it during a drill and reaches a sensible outcome. Runbooks validated only by their author describe what the author already knows.
05When should it be updated?
After every incident, with what was actually useful and what was missing. Runbooks that are written once and never revised describe the failures somebody imagined rather than the ones that happen.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.