Checklist · 5 minute read
Prompt Review Checklist: What to Check Before Merging
A prompt change is a behaviour change and deserves review. Check for conflicting instructions, unsafe handling of injected data, an explicit output contract, sensitive content, and evidence from an evaluation run. A prompt merged without that evidence is an untested change to production behaviour.
A prompt change is a behaviour change and deserves the same review as code. This checklist covers what a reviewer should look for, drawn from FISTA Solutions' AI agents engineering practice.
What does a reviewer look for?
Six categories of defect, in order of frequency.
| Category | What it causes |
|---|---|
| Conflicting instructions | Unpredictable behaviour |
| Unclear data boundaries | Injection exposure |
| No output contract | Downstream parsing failures |
| Sensitive content | Data leakage |
| Accumulated cruft | Prompt nobody dares change |
| No evaluation evidence | Untested behaviour change |
Clarity and consistency
Read the whole prompt as the model receives it, not only the diff.
- The task is stated once, clearly, near the start
- No two instructions contradict each other
- Instructions added for past incidents still apply
- Terminology is consistent throughout
- Negative instructions are minimal and specific
- The prompt reads coherently end to end, not as accumulated layers
- Anything that could be code rather than instruction has been moved
Data boundaries and injection
Retrieved content and user input are data. See why agent security is different.
- Retrieved content clearly delimited from instructions
- User input clearly delimited and labelled as untrusted
- No instruction tells the model to follow directions found in content
- Injection attempts included in the evaluation suite
- Authorisation is enforced in code, not requested in the prompt
- Tool descriptions reviewed as prompts in their own right
- Behaviour tested with adversarial content in the retrieved passages
Output contract
Specify the shape and validate it. See the return of determinism.
- Expected output format stated explicitly
- A schema or structured output mode used where available
- Validation implemented in code for every field
- Behaviour defined when validation fails
- Field names and types match what downstream code expects
- Example outputs included where they help, minimal where they do not
- Length constraints stated where long output is not wanted
Sensitive content
Prompts are drafted against real data and frequently keep it.
- No customer names, identifiers, or account numbers
- No internal system names or credentials
- No API keys, tokens, or connection strings
- Examples are synthetic or properly anonymised
- Nothing that would be a problem if the prompt leaked
- Internal policy detail included deliberately, not accidentally
- The prompt is safe to include in a support ticket
Versioning and provenance
Prompts outside version control are outside change control. See how to set up AI change control.
- The prompt lives in version control alongside the code
- The change is in a pull request with a description of intent
- The prompt version is logged with every output it produces
- Deployment goes through the pipeline, not a vendor console
- The prompt has a named owner
- Rollback to the previous version is a deployment, not an edit
- Related prompts that might conflict have been checked
Evaluation evidence
Without this the reviewer is assessing writing, not behaviour.
- The evaluation suite has been run against the new version
- Results compared against the current production version
- Cases that regressed identified and explained
- The case that motivated the change is now covered by a test
- Cost impact measured if the prompt grew
- Latency impact measured if the prompt grew substantially
- The results are attached to the pull request
What are the most common failures?
Reviewing the diff rather than the whole prompt. Instructions added for incidents that no longer occur. Format requested without validation. Real customer data in examples. And merging without an evaluation run.
Who should own this?
The team that owns the feature owns its prompts, with review by someone who did not write the change. Prompts owned by nobody accumulate instructions until nobody will touch them.
How often should it run?
On every change, plus a periodic full read of production prompts â quarterly is reasonable â to prune accumulated instructions that no longer apply.
What evidence should it produce?
The pull request with its evaluation results, the prompt version history, and a record of which version produced which outputs. That is what makes a behaviour question answerable later.
What about prompts inside vendor products?
You usually cannot see them, which is a real limitation. The questions to ask a vendor are whether prompts are versioned, whether changes are evaluated, and whether you are notified when behaviour changes.
Where you can supply your own instructions or examples, apply this checklist to those. See AI third party risk checklist.
What should you do first?
Read your longest production prompt end to end. Most teams find at least one instruction that contradicts another or addresses a problem that no longer exists.
How FISTA Solutions helps
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: prompts reviewed as behaviour changes with evaluation evidence attached, and versions logged against outputs so quality questions can be traced, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To adapt this checklist to your environment, message FISTA on WhatsApp, or read how to set up AI change control.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What is the most common prompt defect?
Conflicting instructions accumulated over time â one section asking for brevity, another added later asking for detail. The model resolves the conflict unpredictably and the behaviour looks random.
02What should a reviewer check about injected data?
That retrieved content and user input are clearly delimited and described as data, and that no instruction tells the model to follow directions found inside them.
03What is an output contract?
An explicit statement of the expected output shape, with validation in code that enforces it. A prompt that requests a format without validation will occasionally produce something else.
04Why check for sensitive content?
Because prompts are drafted against real examples, and customer names, identifiers, and internal details routinely survive into the committed version.
05What evidence should accompany a change?
An evaluation run comparing the new version against the current one across the suite. Without it, the reviewer is assessing prose rather than behaviour.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.