Checklist ┬╖ 5 minute read
AI Code Review Checklist: Reviewing What You Did Not Write
Generated code is plausible, which is what makes it dangerous to skim. Check it against the specification, look hard at error handling and edge cases, verify security-relevant logic line by line, question every new dependency, and confirm the tests test behaviour rather than restating the implementation.
Generated code is plausible, which is exactly why it needs review rather than a skim. This checklist covers what to check, drawn from FISTA Solutions' AI enablement engineering practice.
What is different about reviewing generated code?
Five differences that change how you read it.
| Difference | Review implication |
|---|---|
| Reads confidently regardless | Confidence is not a signal |
| May solve a nearby problem | Check against the spec |
| No author to ask why | Reasoning must be inferred |
| Adds dependencies freely | Question every import |
| Tests may mirror the code | Read tests against requirements |
| Volume is higher | Review capacity is the constraint |
Correctness against the specification
Read the requirement first, then the code. See spec-driven development in practice.
- A written specification exists for what this should do
- Every requirement in the spec is addressed
- Nothing is implemented that the spec did not ask for
- Assumptions the code makes are valid in your system
- Business rules match the actual rules, not plausible ones
- Calculations verified independently, not read for plausibility
- Naming reflects what the code does, not what it was asked to do
Error handling and edge cases
The happy path is usually right; this is where the defects concentrate.
- Empty, null, and missing inputs handled
- Boundary values handled at both ends
- External call failures handled explicitly
- Timeouts set on anything that leaves the process
- Errors surfaced rather than swallowed
- Partial failure leaves recoverable state
- Concurrency and ordering assumptions checked
Security
Read this line by line. Plausibility is not a defence.
- Input validated before use
- Queries parameterised, never concatenated
- Authorisation checked, not assumed from context
- No credentials or secrets in the code
- Output encoded appropriately for its destination
- Logging does not record sensitive values
- Cryptographic code uses standard libraries, not hand-rolled logic
Dependencies
Every import is a long-term commitment made by something that will not maintain it.
- Each new dependency is justified
- The codebase does not already solve this
- The library is maintained and widely used
- Licence is compatible with your use
- Version is pinned
- The dependency actually exists and is the one intended
- Transitive additions reviewed, not just the direct one
Tests
Read tests against the requirement, not against the implementation.
- Tests assert required behaviour, not implementation detail
- Edge cases and error paths covered, not only the happy path
- A test would fail if the implementation were wrong in a plausible way
- No assertions that merely restate what the code does
- Test data is realistic
- Tests are deterministic and do not depend on timing
- Coverage of the specific requirement, not just line coverage
Codebase fit
Generated code follows general conventions, not yours.
- Existing utilities used rather than reimplemented
- Project conventions followed for structure and naming
- Error handling matches the codebase's approach
- Logging matches the codebase's format and levels
- No duplicated logic that exists elsewhere
- Abstractions consistent with surrounding code
- Comments explain why, not what
What are the most common failures?
Skimming because it reads well. Reviewing the code without the requirement. Accepting tests that mirror the implementation. Waving through dependencies. And approving volume because the queue is long.
Who should own this?
The reviewer owns the merged result as much as the author would. That is the point: accountability does not transfer to the tool, and the reviewer is the accountable party.
How often should it run?
Every change. Additionally, a periodic audit of generated code merged during busy periods, because review quality falls under queue pressure and the defects stay.
What evidence should it produce?
Review comments showing the specification was consulted, and test results. Over time, defect rates in generated versus hand-written code tell you whether review is working.
What if the volume is unreviewable?
Then that is the finding, and it is a real one. A team generating more code than it can review has moved the bottleneck rather than removed it.
The responses are smaller changes, stricter merge criteria, and treating review capacity as the planning constraint rather than generation capacity. See the new shape of engineering teams.
What should you do first?
Pick a recently merged generated change and review it properly against its requirement. What you find tells you whether your current review is working.
How FISTA Solutions helps
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: generated code reviewed against a written specification rather than for plausibility, with review capacity treated as the planning constraint, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To adapt this checklist to your environment, message FISTA on WhatsApp, or read spec-driven development in practice.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why is generated code harder to review?
Because it reads well. Human-written code with a misunderstanding often looks confused; generated code with a misunderstanding looks confident, which defeats the instinct reviewers rely on.
02What should you check first?
Whether it does what the specification asked. Generated code frequently solves a nearby problem correctly, and that is invisible unless you compare against the requirement rather than reading the code.
03Where are the common defects?
Error handling, boundary conditions, concurrency, and anything requiring knowledge of your specific system. The happy path is usually correct; the edges are where it fails.
04Why scrutinise dependencies?
Because generated code readily imports libraries to solve small problems, adding maintenance and supply chain surface for something the codebase could already do.
05What is wrong with generated tests?
They frequently assert what the implementation does rather than what it should do, so a bug in the implementation produces a passing test that codifies the bug.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.