Playbook ┬╖ 6 minute read
How to Run an AI Incident Postmortem That Changes Things
An AI incident postmortem works when the trajectory can be reconstructed, the analysis goes past the prompt to why the system was allowed to be wrong unchecked, and the actions become owned changes with dates rather than recommendations in a document nobody reopens.
AI incidents are harder to reconstruct than outages, because nothing failed тАФ the system produced a confident wrong result and everything downstream worked as designed. This playbook covers postmortems that produce changes, drawing on FISTA Solutions' AI agents work.
When is this worth doing?
After any incident where an AI system produced a materially wrong output that reached a person, took an action it should not have, or failed in a way that required intervention.
Also after significant near-misses. A wrong output caught by a reviewer is the same failure with a better outcome, and analysing it is considerably cheaper than waiting for the version that was not caught.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Reconstruct the trajectory | Inputs, context, steps, output |
| 2. Establish the timeline | Detection, escalation, resolution |
| 3. Analyse causes | Past the prompt, to the controls |
| 4. Ask why nothing caught it | The more productive question |
| 5. Agree actions | Owners and dates attached |
| 6. Add the test case | So it cannot recur silently |
Step 1 тАФ Reconstruct the trajectory
Assemble what the system received, what it retrieved, what steps it took, and what it produced.
This is only possible if it was logged. Systems that record outputs without inputs, or final answers without retrieved context, cannot be analysed after the fact, and that gap is itself the first finding of the postmortem.
Where logs are incomplete, say so explicitly in the writeup rather than reasoning around the gap. Speculation recorded as analysis leads to fixes for the wrong problem.
Step 2 тАФ Establish the timeline
When did the system produce the wrong result, when was it noticed, by whom, how did it escalate, and when was it resolved.
The gap between occurrence and detection is usually the most informative number. An incident detected by a customer three weeks after it started reveals a monitoring problem larger than whatever caused the original error.
Record how it was detected too. If the answer is always 'a customer complained', that is a standing finding rather than an incident-specific one.
Step 3 тАФ Analyse causes past the prompt
The prompt is the easiest thing to blame and rarely the real cause.
Look further: was the retrieved context wrong or stale, did the evaluation set cover this case, did the system have authority it should not have had, had the model version changed, was the input type one nobody anticipated.
Fixing a prompt in response to an incident feels productive and usually addresses one instance of a class. The class is what matters, and finding it requires asking what else could produce the same outcome.
Step 4 тАФ Ask why nothing caught it
A wrong output is expected occasionally. A wrong output reaching a person unchecked is a control failure, and that is the more productive analysis.
Was there a confidence threshold that should have triggered abstention, a human review step that was bypassed or absent, a monitor that should have alerted, an escalation path that was not taken. Each of those is a control, and each can be strengthened for a whole class of failures rather than one.
This question also moves the conversation away from the model's fallibility, which is a fixed property, towards the system's design, which is not.
Step 5 тАФ Agree actions with owners and dates
Every action needs a named person and a date, recorded somewhere that gets reviewed.
Actions without both are recommendations. The reliable signal of a postmortem culture that is not working is the same action appearing in three consecutive writeups.
Keep the list short. Three actions that get done beat eleven that get filed, and the discipline of choosing produces better prioritisation than the appearance of thoroughness.
Step 6 тАФ Add the test case
Every incident becomes a case in the evaluation suite, with the input that caused it and the expected correct behaviour.
That single practice is what makes evaluation progressively better at catching real problems rather than imagined ones. It also means the specific failure cannot recur silently, which is the minimum a postmortem should guarantee.
If the input contains customer data, construct an equivalent case rather than skipping it. See how to run an ai evaluation program.
How do you keep it blameless in practice?
By writing the process down, running the first few carefully, and having a senior person go first when something they own is involved.
Blameless is a practice rather than a declaration. The test is whether someone can say they approved a change that turned out badly without it affecting how they are treated, and people learn the answer by watching rather than by reading the policy.
What about incidents involving customers or regulators?
They carry notification obligations that run on their own timelines, and those take priority over the analysis.
Run the notification process and the postmortem in parallel with different people where possible. Analysis conducted under notification deadline pressure is rushed, and the resulting actions are correspondingly shallow. See how to handle an ai data breach.
Who needs to be involved?
Whoever operates the system, whoever built it, someone from the function affected, and a facilitator who did not work on it.
The outside facilitator matters more than it seems. Teams analysing their own incident stop at the explanation they already believed.
How long does it take?
A focused session of one to two hours within a few days of resolution, plus writing up. Postmortems held weeks later analyse people's memories rather than the incident.
What are the common failure modes?
No trajectory logs to reconstruct from. Stopping at the prompt. Not asking why nothing caught it. Actions without owners. Long action lists nobody completes. And no test case added.
How do you know it worked?
Actions completed by their dates, the specific failure represented in the evaluation suite, detection time falling across incidents, and people volunteering near-misses rather than hiding them.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
Check whether you could reconstruct your last AI incident from logs alone. If not, adding trajectory logging is the action that makes every future postmortem possible.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: trajectory logging in place so incidents can be reconstructed, every incident converted into an evaluation case with an owner and a date, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read how to run an AI evaluation program.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why are AI incidents hard to reconstruct?
Because the system did not fail тАФ it produced a confident wrong result. Without logged inputs, retrieved context, and intermediate steps, there is no way to establish what happened, and those logs cannot be created afterwards.
02What is usually the real cause?
Rarely the prompt. More often it is missing evaluation coverage, a knowledge base containing something stale or contradictory, authority the system should not have had, or a monitoring gap that let it continue.
03What question matters most?
Why nothing caught it. A wrong output is expected occasionally; a wrong output reaching a customer unchecked is a control failure, and that is the more productive line of analysis.
04Why blameless?
Because the alternative produces incomplete information. People who expect blame describe what they did defensively, and the postmortem then analyses a sanitised version of events that fixes nothing.
05What makes actions actually happen?
Named owners, dates, and a scheduled review of whether they were completed. Actions recorded without all three are recommendations rather than commitments, and the reliable sign of a postmortem practice that is not working is the same action appearing in three consecutive writeups.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.