Playbook ┬╖ 6 minute read
How to Debug an AI Agent When It Goes Wrong
Debugging an agent requires reproducing the trajectory, isolating which step failed, and distinguishing planning failures from tool failures from context failures. Without recorded trajectories none of that is possible, which makes observability the prerequisite for the whole exercise rather than an optional addition.
Debugging an agent is harder than debugging a conventional service because the control flow is decided at runtime by a non-deterministic component. This playbook covers doing it systematically, drawing on FISTA Solutions' AI agents work.
When is this worth doing?
Whenever an agent produces a wrong result, takes an unexpected action, or behaves inconsistently тАФ and before anyone adjusts the prompt in response.
The instinct to change the prompt first is strong and usually wrong. Prompt changes without diagnosis fix one instance of a class and frequently break something else.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Capture the trajectory | Input, steps, tool calls, output |
| 2. Reproduce it | Same inputs, not similar ones |
| 3. Isolate the failing step | The first wrong output, not the symptom |
| 4. Classify the failure | Planning, tool, context, or output |
| 5. Fix the right layer | Each class has a different fix |
| 6. Add the regression test | Because agent fixes regress |
Step 1 тАФ Capture the trajectory
You need the input, the model's stated reasoning where the provider exposes it, every tool call with its arguments and its result, any retrieved context, and the final output.
Systems that log only the final output cannot be debugged. This is the single most common gap, and it converts every investigation into speculation.
If the trajectory was not recorded, the first action is adding the logging and waiting for a recurrence. That is frustrating and faster than guessing. See what is distributed tracing for ai.
Step 2 тАФ Reproduce with the same inputs
Replay the exact trajectory: the same input, the same retrieved context, the same model version, the same tool responses.
Similar inputs are not enough. Agents are sensitive to phrasing, and a paraphrased version of the failing question frequently works, which sends the investigation in the wrong direction and produces a conclusion that the problem was intermittent.
Where tool responses depended on state that has changed, replay with the recorded responses rather than live ones. Otherwise you are debugging a different situation.
Step 3 тАФ Isolate the first wrong step
Walk the trajectory and find the first step whose output is wrong given its input.
That is usually earlier than where the symptom appeared. An agent that produced a wrong final answer may have gone wrong three steps earlier by calling the wrong tool, and everything after that was a reasonable response to bad information.
Being precise here is what prevents fixing the wrong thing. The final step is where the problem is visible and rarely where it started.
Step 4 тАФ Classify the failure
Planning failures: the agent chose an approach that could not work, or called the wrong tool for the task. Tool failures: a call returned an error, empty result, or unexpected format the agent handled badly. Context failures: necessary information was absent, truncated, or superseded by more recent content. Output failures: the process was right and the final answer was not.
Each has a different fix, and treating them all as prompt problems is why agents get progressively more complicated prompts without getting better.
Write the classification down. Patterns across incidents are more informative than any single case.
Step 5 тАФ Fix the layer that failed
Planning failures respond to clearer tool descriptions, fewer tools, or decomposing the task. Tool failures respond to better error handling, clearer result formats, and validation. Context failures respond to retrieval work, summarisation, or larger budgets. Output failures respond to prompting and formatting constraints.
Prefer structural fixes to instructional ones. Removing a tool the agent keeps misusing is more reliable than instructing it not to, and restricting what a tool accepts is more reliable than asking for well-formed arguments.
Change one thing and re-run. Bundled changes make attribution impossible for the next failure.
Step 6 тАФ Add the regression test
Every fixed failure becomes a case in the evaluation suite with the same input and the expected behaviour.
Agent fixes regress unusually easily. Prompt changes have effects beyond the case they targeted, tool changes affect every path using that tool, and model updates change everything at once.
Without the test, the same bug returns and gets fixed again, usually by someone who does not know it was fixed before. See what is a regression suite for ai.
How do you debug intermittent failures?
By looking at what varies between runs, which is usually context rather than code.
Retrieved documents differ as the corpus changes. Conversation history accumulates until truncation drops something important. Tool responses vary with system state. Temperature settings introduce variation in the model itself.
Collect several trajectories of the same task тАФ successful and failed тАФ and diff them. The difference is almost always in the context, and it is usually visible immediately once you look.
What about failures after a model update?
Suspect the update first, and check whether the model version changed.
Systems calling a provider without pinning a version can change behaviour without any deployment, which produces failures nobody can attribute to a change. That is a strong argument for pinning versions, quite apart from the compliance reasons.
If a version change is the cause, the fix is usually prompt adjustment plus re-running the full evaluation suite rather than patching the specific failure.
Who needs to be involved?
An engineer with access to trajectories, and someone who can say what the correct behaviour would have been.
The second role matters more than it appears. Debugging towards the wrong expected outcome is a productive-feeling way to waste a day.
How long does it take?
Hours for a well-instrumented system with a reproducible case. Days or longer if trajectories were not recorded, most of which is waiting for a recurrence with logging in place.
What are the common failure modes?
Changing the prompt before diagnosing. No recorded trajectories. Reproducing with similar rather than identical inputs. Fixing the last step instead of the first wrong one. And bundling changes.
How do you know it worked?
Failures diagnosed to a layer rather than guessed at, fixes that do not regress, and a growing regression suite that catches recurrences before users do.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
Check whether you could replay your last agent failure from logs. If not, adding trajectory recording is the highest-value change available.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: trajectory recording in place so failures can be replayed rather than guessed at, fixes applied to the layer that actually failed, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read hire agent engineers.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What do you need before you can debug an agent?
Recorded trajectories: the input, the model's reasoning where available, every tool call with arguments and results, and the final output. Without them you are reasoning about what might have happened rather than what did.
02What are the main failure classes?
Planning failures where the agent chose the wrong approach, tool failures where a call returned something unexpected, context failures where necessary information was missing or lost, and output failures where a correct process produced a bad final answer.
03How do you isolate the failing step?
Replay the trajectory and examine each step's inputs and outputs in order. The failure is at the first step where the output is wrong given its input, which is frequently earlier than where the symptom appeared.
04Why are intermittent failures usually context?
Because context varies between runs in ways code does not: retrieved documents differ, conversation history accumulates, and truncation drops different material. Those produce failures that reproduce only under specific conditions.
05What happens after the fix?
The failing case becomes a regression test with the same input and expected behaviour. Agent fixes regress easily because changes to prompts and tools have effects beyond the case being fixed.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.