FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Playbook ¡ 6 minute read

How to Run an AI Incident Postmortem That Changes Things

An AI incident postmortem works when the trajectory can be reconstructed, the analysis goes past the prompt to why the system was allowed to be wrong unchecked, and the actions become owned changes with dates rather than recommendations in a document nobody reopens.

By FISTA Solutions¡ AI-Native Engineering Team¡
How to Run an AI Incident Postmortem That Changes Things article cover

AI incidents are harder to reconstruct than outages, because nothing failed — the system produced a confident wrong result and everything downstream worked as designed. This playbook covers postmortems that produce changes, drawing on FISTA Solutions' AI agents work.

When is this worth doing?

After any incident where an AI system produced a materially wrong output that reached a person, took an action it should not have, or failed in a way that required intervention.

Also after significant near-misses. A wrong output caught by a reviewer is the same failure with a better outcome, and analysing it is considerably cheaper than waiting for the version that was not caught.

What does the sequence look like?

StepPurpose
1. Reconstruct the trajectoryInputs, context, steps, output
2. Establish the timelineDetection, escalation, resolution
3. Analyse causesPast the prompt, to the controls
4. Ask why nothing caught itThe more productive question
5. Agree actionsOwners and dates attached
6. Add the test caseSo it cannot recur silently

Step 1 — Reconstruct the trajectory

Assemble what the system received, what it retrieved, what steps it took, and what it produced.

This is only possible if it was logged. Systems that record outputs without inputs, or final answers without retrieved context, cannot be analysed after the fact, and that gap is itself the first finding of the postmortem.

Where logs are incomplete, say so explicitly in the writeup rather than reasoning around the gap. Speculation recorded as analysis leads to fixes for the wrong problem.

Step 2 — Establish the timeline

When did the system produce the wrong result, when was it noticed, by whom, how did it escalate, and when was it resolved.

The gap between occurrence and detection is usually the most informative number. An incident detected by a customer three weeks after it started reveals a monitoring problem larger than whatever caused the original error.

Record how it was detected too. If the answer is always 'a customer complained', that is a standing finding rather than an incident-specific one.

Step 3 — Analyse causes past the prompt

The prompt is the easiest thing to blame and rarely the real cause.

Look further: was the retrieved context wrong or stale, did the evaluation set cover this case, did the system have authority it should not have had, had the model version changed, was the input type one nobody anticipated.

Fixing a prompt in response to an incident feels productive and usually addresses one instance of a class. The class is what matters, and finding it requires asking what else could produce the same outcome.

Step 4 — Ask why nothing caught it

A wrong output is expected occasionally. A wrong output reaching a person unchecked is a control failure, and that is the more productive analysis.

Was there a confidence threshold that should have triggered abstention, a human review step that was bypassed or absent, a monitor that should have alerted, an escalation path that was not taken. Each of those is a control, and each can be strengthened for a whole class of failures rather than one.

This question also moves the conversation away from the model's fallibility, which is a fixed property, towards the system's design, which is not.

Step 5 — Agree actions with owners and dates

Every action needs a named person and a date, recorded somewhere that gets reviewed.

Actions without both are recommendations. The reliable signal of a postmortem culture that is not working is the same action appearing in three consecutive writeups.

Keep the list short. Three actions that get done beat eleven that get filed, and the discipline of choosing produces better prioritisation than the appearance of thoroughness.

Step 6 — Add the test case

Every incident becomes a case in the evaluation suite, with the input that caused it and the expected correct behaviour.

That single practice is what makes evaluation progressively better at catching real problems rather than imagined ones. It also means the specific failure cannot recur silently, which is the minimum a postmortem should guarantee.

If the input contains customer data, construct an equivalent case rather than skipping it. See how to run an ai evaluation program.

How do you keep it blameless in practice?

By writing the process down, running the first few carefully, and having a senior person go first when something they own is involved.

Blameless is a practice rather than a declaration. The test is whether someone can say they approved a change that turned out badly without it affecting how they are treated, and people learn the answer by watching rather than by reading the policy.

What about incidents involving customers or regulators?

They carry notification obligations that run on their own timelines, and those take priority over the analysis.

Run the notification process and the postmortem in parallel with different people where possible. Analysis conducted under notification deadline pressure is rushed, and the resulting actions are correspondingly shallow. See how to handle an ai data breach.

Who needs to be involved?

Whoever operates the system, whoever built it, someone from the function affected, and a facilitator who did not work on it.

The outside facilitator matters more than it seems. Teams analysing their own incident stop at the explanation they already believed.

How long does it take?

A focused session of one to two hours within a few days of resolution, plus writing up. Postmortems held weeks later analyse people's memories rather than the incident.

What are the common failure modes?

No trajectory logs to reconstruct from. Stopping at the prompt. Not asking why nothing caught it. Actions without owners. Long action lists nobody completes. And no test case added.

How do you know it worked?

Actions completed by their dates, the specific failure represented in the evaluation suite, detection time falling across incidents, and people volunteering near-misses rather than hiding them.

What does it cost?

Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.

Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.

What should you do first?

Check whether you could reconstruct your last AI incident from logs alone. If not, adding trajectory logging is the action that makes every future postmortem possible.

How FISTA Solutions helps

FISTA Solutions runs this work alongside client teams rather than around them: trajectory logging in place so incidents can be reconstructed, every incident converted into an evaluation case with an owner and a date, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.

To run this with support, message FISTA on WhatsApp, or read how to run an AI evaluation program.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01Why are AI incidents hard to reconstruct?

Because the system did not fail — it produced a confident wrong result. Without logged inputs, retrieved context, and intermediate steps, there is no way to establish what happened, and those logs cannot be created afterwards.

02What is usually the real cause?

Rarely the prompt. More often it is missing evaluation coverage, a knowledge base containing something stale or contradictory, authority the system should not have had, or a monitoring gap that let it continue.

03What question matters most?

Why nothing caught it. A wrong output is expected occasionally; a wrong output reaching a customer unchecked is a control failure, and that is the more productive line of analysis.

04Why blameless?

Because the alternative produces incomplete information. People who expect blame describe what they did defensively, and the postmortem then analyses a sanitised version of events that fixes nothing.

05What makes actions actually happen?

Named owners, dates, and a scheduled review of whether they were completed. Actions recorded without all three are recommendations rather than commitments, and the reliable sign of a postmortem practice that is not working is the same action appearing in three consecutive writeups.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project