Glossary ¡ 5 minute read
What Is Agent Reflection? Self-Critique in AI Systems Explained
Agent reflection is the step in which a system reviews its own output or plan against criteria before finalising it. It reliably catches format errors, omissions, and constraint violations. It does not reliably catch factual errors, because the same model that produced the mistake is judging it, which is why external verification matters more.
Reflection is one of the most widely adopted agent patterns and one of the most widely over-credited. It genuinely improves outputs in specific, identifiable ways, and it cannot do the thing teams most want from it, which is catching the model's own factual mistakes. This explainer covers the distinction. It complements what is an agent loop and how to build an agent evaluation harness, and reflects FISTA Solutions' approach in AI agents delivery.
What is reflection?
A step after generation in which the model is shown its own output and asked to evaluate it, typically against criteria, and to revise if needed. In agent systems it may review a plan before execution or a result before returning it.
The pattern is simple to implement, which partly explains its popularity, and its effectiveness varies enormously by what it is asked to check.
| Error type | Reflection catches | Better method |
|---|---|---|
| Missing required section | Reliably | â |
| Wrong output format | Reliably | Schema validation |
| Explicit constraint violated | Reliably | Rule check in code |
| Skipped step against checklist | Reliably | â |
| Factual error | Poorly | Source verification |
| Faulty reasoning | Poorly | Different model or tool |
What does it catch well?
Structural problems. Did the output include every required section, does it follow the specified format, does it satisfy the stated constraints, were any checklist items skipped. These are verifiable by inspection against explicit criteria, and a model does this genuinely well.
Much of that could also be checked in code, which is cheaper and deterministic. Reflection earns its place where the criteria require judgement â is this tone appropriate, does this actually address the question asked.
Why does it miss factual errors?
Because the reviewer is the writer. The same model, with the same training and the same gaps, evaluates a claim it produced confidently. A confident wrong statement looks correct on review, because the reason it was produced is the reason it is believed.
This is a structural limitation, not a prompting failure. Asking the model to be more critical produces more critical-sounding output, not more accurate judgement, and sometimes causes it to change correct answers.
What works better for facts?
External verification. Retrieving the source document and comparing the claim against it. Executing the code to see whether it runs. Querying the system of record. Using a different model with different training data as the reviewer.
The common property is introducing information that was not present when the output was generated. Reflection that includes a retrieval step is substantially more useful than reflection over the output alone. See what is retrieval augmentation.
How should criteria be specified?
Explicitly, as a checklist rather than an invitation. "Review this and improve it" produces changes of uncertain value. "Check that each of these five requirements is met, and state which are not" produces a verifiable result that can be acted on and measured.
Explicit criteria also make the reflection step evaluable: you can measure how often it correctly identifies a known-missing element.
How many passes are worth running?
Usually one. A second occasionally catches something, and beyond that models tend to make changes for their own sake, sometimes degrading a correct answer into a worse one. Unbounded reflection loops are a known failure mode and produce both cost and drift.
The iteration limit belongs in application code, not in the prompt.
What does it cost?
A full model call per pass with the original output in context, plus any regeneration triggered. On high-volume workloads that can approximately double cost, which is significant enough that reflection should be applied selectively â to high-stakes outputs, or where measurement shows the error rate justifies it.
Where does it fit with human review?
Before it, narrowing what a human must examine. Reflection that flags "requirement three is not addressed" makes human review faster and more focused. Reflection presented as an assurance that output is correct makes human review less careful, which is the opposite of the intended effect.
How the result is framed to the reviewer therefore matters as much as the reflection itself.
What should you do first?
Categorise the errors your system actually makes. If they are structural, reflection or code validation will help immediately. If they are factual, reflection will not, and the effort belongs in retrieval and verification instead. That single categorisation prevents the most common misapplication of the pattern.
How does it differ from a critic model?
A critic is a separate model, often smaller and specifically trained or prompted to judge, applied to another model's output. Because it brings different parameters and sometimes different training data, it does not share the generator's blind spots in the way self-reflection does.
That makes critics more useful for factual and reasoning checks and more expensive to operate, since they must be maintained and evaluated as their own component. For most teams, external verification against a source is the cheaper route to the same benefit.
How FISTA Solutions helps
FISTA Solutions applies reflection against explicit criteria for structural checks, uses external verification and tool execution for factual claims, bounds reflection iterations in application code, applies it selectively where error rates justify the cost, and frames results to sharpen rather than substitute for human review, through AI agents, AI enablement, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To improve agent output quality where it can actually be improved, message FISTA on WhatsApp, or read what is an agent loop.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What does reflection reliably catch?
Structural problems. Missing required sections, wrong output format, unmet explicit constraints, and steps skipped against a stated checklist. These are verifiable against criteria, and a model checking its own output against a list performs this genuinely well.
02Why does it miss factual errors?
Because the same model that produced the error is evaluating it, with the same knowledge and the same blind spots. A confidently wrong claim looks correct on review, which is why self-critique is not a substitute for checking facts against a source.
03What works better for factual checking?
External verification: retrieving the source and comparing, running the code, querying the system of record, or having a different model with different training review it. Anything that introduces information the original generation did not have.
04How many reflection passes are useful?
Usually one. A second pass occasionally catches something the first missed, and beyond that the model tends to make changes for the sake of changing, sometimes degrading a correct answer. Iteration counts should be bounded explicitly.
05What does reflection cost?
Another full model call per pass, with the original output in context, plus any regeneration it triggers. On high-volume workloads that can double cost, so it should be applied where the error rate justifies it rather than universally.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.