Playbook · 6 minute read
How to Build a Test Generation Agent for Software Teams
A test generation agent writes tests against intended behaviour taken from specifications, issues, and interfaces rather than from the implementation, verifies test quality by whether the tests fail when the code is deliberately broken, and routes every generated test through human review. Coverage without defect detection is theatre.
Test generation is among the easiest things to demonstrate with a coding model and among the easiest to get wrong in a way that looks successful. Coverage rises, the suite grows, and the tests catch nothing, because they were written from the implementation and assert that the code does what it already does. Getting value requires discipline about where tests come from and how their quality is measured. This guide covers building it properly, drawing on FISTA Solutions' AI agents work in engineering. It complements ai test generation and ai code review. This article is general guidance, not legal advice.
What is the fundamental trap?
Tests derived from the implementation. A model shown a function and asked to test it will produce assertions matching its current behaviour, including any defect. The resulting test passes permanently, raises the coverage number, and would not detect the bug it has just canonised.
Worse, when someone later fixes the bug, the test fails and appears to indicate a regression. Implementation-derived tests actively obstruct correctness.
| Test source | Detection value | Notes |
|---|---|---|
| Implementation code | Near zero | Asserts current behaviour |
| Specification or issue | High | States intent |
| Interface contract | High | Defines obligations |
| Bug report | Very high | Known failure case |
| Property definition | High | Invariants across inputs |
| Existing test patterns | Moderate | Style only, not assertions |
Where should tests come from?
Stated intent. Specifications, acceptance criteria, issue descriptions, interface contracts, API schemas, and bug reports all describe what the code should do independently of what it does. Tests generated from those sources have genuine detection value.
Where intent is not documented — which is common — the honest path is generating a proposed specification from the code for a developer to confirm or correct, and then generating tests from the confirmed specification. That two-step flow surfaces disagreements about intended behaviour, which is valuable in itself.
How do you measure whether the tests are any good?
Mutation testing. Introduce deliberate faults into the code — invert a condition, change a boundary, remove a statement — and check whether the test suite fails. A suite that passes against broken code provides no protection, whatever its coverage percentage.
This is the only honest measurement, and it should gate the agent's output: generated tests that do not kill mutants do not get merged. Building the mutation harness alongside the generator, rather than afterwards, is what keeps the programme from producing coverage theatre.
Where does generation add real value?
Edge cases. Empty collections, single elements, maximum sizes, null and undefined handling, unicode, timezone and daylight-saving boundaries, concurrent access, and error paths. Developers know these matter and skip them under deadline pressure, consistently.
Systematic enumeration of boundaries is something a model does better than a tired human at 6pm, and the defects found are real. This is the strongest argument for the whole approach.
Why does review stay mandatory?
Because a wrong test is worse than no test. It encodes incorrect behaviour as a requirement, blocks the correct fix later, and carries the authority of being in the suite. Someone must confirm that each assertion reflects intended behaviour.
This constrains volume. Generating two hundred tests nobody can meaningfully review produces rubber-stamping, which is the failure the control exists to prevent. Fewer, higher-value tests that get genuine review beat comprehensive suites that get skimmed. See ai code review.
What about flaky tests?
They must be caught before merge by running each generated test repeatedly — twenty or more times — and rejecting anything non-deterministic. Timing dependencies, ordering assumptions, and shared state are common in generated tests because the model cannot see the runtime environment.
A flaky test in a shared suite is a long-lived tax: it teaches everyone to re-run failing builds rather than investigate, and it eventually hides a real failure.
How should integration and end-to-end tests be treated?
Carefully, and mostly not by generation. They depend on environment, data setup, and system behaviour a model cannot observe, and the failure modes are slow and expensive. Unit and component tests are where generation performs well; broader tests need human design with generation assisting on data setup and assertions.
What about tests for existing untested code?
The most requested use and the most dangerous, because intent is undocumented and the code is the only source. The workable approach is generating a behaviour description for developer confirmation first, correcting it where the current behaviour is wrong, and only then generating tests. Skipping that step encodes legacy bugs permanently.
How does it integrate?
Into the pull request workflow, proposing tests alongside a change, with mutation results and flakiness checks visible in the review. Tests generated in a separate tool and pasted in lose the context that makes them reviewable.
How is it evaluated?
On mutation score, defects caught before production, review time per generated test, flaky test rate, and suite runtime growth. Coverage percentage is the metric that rises fastest and means least, and it is the one most likely to be reported.
What does the build sequence look like?
One week on the mutation testing harness, first, because it is the quality gate. Two weeks on intent-sourced generation for a single well-specified module. One week on flakiness screening. One week on pull request integration. Expand to further modules only once the mutation score holds.
What goes wrong?
Generation from implementation. Coverage as the success metric. Volume that defeats review. No mutation gate. Flaky tests merged. End-to-end generation. And retrofitting tests onto legacy code without confirming intent.
What does it cost to run?
Generation is inexpensive; mutation testing is the significant compute cost because it runs the suite many times. Scoping mutation runs to changed code rather than the whole repository keeps it proportionate while preserving the signal where it matters.
What does good look like after six months?
Mutation score rising rather than coverage, edge-case defects being caught in review instead of production, generated tests reviewed as carefully as handwritten ones because there are few enough to read, and no measurable growth in flaky failures.
How FISTA Solutions helps
FISTA Solutions builds test generation into engineering workflows with intent-sourced generation, mutation testing as the quality gate, flakiness screening before merge, pull request integration, and deliberate volume limits that keep review meaningful, through AI agents, AI enablement, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 47% efficiency gains.
To generate tests that actually catch defects, message FISTA on WhatsApp, or read ai code review.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why is implementation-derived testing the core trap?
Because a test written from the code asserts that the code does what it does, including its bugs. It passes forever, adds coverage, and catches nothing. Tests must come from intended behaviour — specifications, issues, interface contracts — to have any detection value.
02How do you know generated tests are any good?
Mutation testing. Deliberately introduce faults into the code and check whether the tests fail. A suite that passes against broken code is worthless regardless of its coverage number, and this is the only measurement that answers the question honestly.
03Where does generation add the most value?
Edge cases and boundaries. Empty inputs, maximum values, null handling, concurrent access, unicode, timezone boundaries — the cases developers know about and skip under deadline pressure. Systematic enumeration is genuinely better than tired human recall.
04Why does review remain necessary?
Because a wrong test is worse than no test: it encodes incorrect behaviour as a requirement and blocks the correct fix later. Generating more tests than anyone can review turns review into rubber-stamping, which is the failure this control exists to prevent.
05What about flaky tests?
They must be caught before merge by running each generated test repeatedly. A flaky test in a shared suite teaches the whole team to ignore failures, and the cost lands on every engineer for years. This is general guidance, not legal advice.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.