Playbook · 6 minute read
How to Run an AI Evaluation Program That Catches Problems
An evaluation programme works when the set reflects real inputs including failures, the rubric produces agreement between reviewers, the suite runs automatically on every change, and someone acts when it regresses. Programmes missing any of those measure activity rather than quality.
Evaluation is the difference between knowing whether an AI system works and believing it does. Most programmes fail for one of two reasons: the set does not reflect reality, or nobody acts on what it finds. This playbook covers avoiding both, drawing on FISTA Solutions' AI enablement work.
When is this worth doing?
Before any AI system reaches production, and immediately for any system already there without it.
There is no threshold at which evaluation becomes worthwhile. A system without it cannot be changed safely, which means it either stops being improved or gets changed on guesswork, and both outcomes are worse than the setup cost.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Collect real inputs | Including the ones that failed |
| 2. Agree what good means | A rubric reviewers agree on |
| 3. Establish the baseline | Where the system is today |
| 4. Automate the run | Every change, not on request |
| 5. Set thresholds and gates | A regression blocks the change |
| 6. Extend from production | Failures become test cases |
Step 1 â Collect real inputs, including failures
Build the set from what the system actually sees: real questions, real documents, real edge cases. Weight it towards the difficult end.
A set of straightforward cases the system handles well produces a high score and no information. The useful cases are the ones near the boundary â ambiguous inputs, unusual formats, questions the system should refuse.
If the system is already in production, mine the complaints and the escalations. Those are pre-labelled failures, and they are the highest-value cases available.
Step 2 â Agree what good means
Write a rubric that states what a correct output looks like, with worked examples of acceptable and unacceptable answers and the reason for each.
Then test it: have two people rate the same twenty outputs independently and compare. Low agreement means the rubric is ambiguous, and no amount of rater training fixes an ambiguous rubric.
Expect two or three iterations. The first version always contains a criterion that seemed obvious and turns out to mean different things to different people. See what is an evaluation rubric.
Step 3 â Establish the baseline
Run the current system against the set and record the result, with the model version, prompt version, and date.
That number is the reference point for everything afterwards. Without it, a later measurement is just a number, and arguments about whether quality has declined become arguments about memory.
Record the failures individually, not just the aggregate. The pattern of what fails is more useful than the score, and it tells you where to work.
Step 4 â Automate the run
Wire the suite into the change process so it runs on every prompt change, model change, retrieval corpus update, and configuration change.
Manual evaluation runs when someone remembers, which is not when it matters. Automated evaluation runs when the risk is created.
The corpus point is easily missed: a retrieval system whose knowledge base changed has effectively changed, and evaluation that only triggers on code changes will not notice. See what is continuous evaluation.
Step 5 â Set thresholds and make them gates
Decide what score is acceptable and what constitutes a regression, and make a regression block the change rather than raise a notification.
A gate is what makes the programme a control rather than a report. Notifications get muted; gates get attention.
Set thresholds per category rather than only in aggregate. A system whose overall score holds steady while failing an entire category of important cases has regressed in a way an average hides.
Step 6 â Extend the set from production
Every failure in production becomes a test case. That single practice is what makes an evaluation suite progressively more useful.
It requires a route from a complaint or escalation to the evaluation set, with someone whose job includes walking it. Without that route, the suite stays as good as the day it was built while production keeps finding new ways to break.
Retire cases too. A test the system has passed for a year on every run is providing little information, and a large slow suite gets run less often.
Who should do the rating?
People who understand the domain, working from the rubric, with agreement measured regularly.
Engineers rating outputs in a domain they do not know produce confident wrong labels, and evaluation built on those is worse than none because it creates false confidence. Domain experts are the right raters and their time is the constraint, which is why the set should stay focused rather than comprehensive. See hire AI trainers.
Can models do the rating?
Partly, and with care. Model-based scoring works reasonably for some criteria â format compliance, presence of required elements, obvious contradictions â and poorly for judgement calls in a specialist domain.
Validate any automated scoring against human ratings before trusting it, and re-validate when models change. A scorer that drifts produces a quality signal that is itself wrong, which is the worst failure available because it looks like data.
Who needs to be involved?
An owner accountable for the programme, domain experts to rate, and an engineer to automate the runs.
The owner matters most. Suites without an owner stop being extended, and a suite that stopped growing two years ago is measuring an old system.
How long does it take?
Two to three weeks to build a useful first set and rubric, a few days to automate, then continuous. Programmes still in design after two months have usually over-scoped the first set.
What are the common failure modes?
Sets built from easy cases. Rubrics nobody tested for agreement. Manual runs. Results with no owner. No gate. And sets that never grow from production failures.
How do you know it worked?
Regressions caught before release rather than by users, agreement between raters holding steady, the set growing from real failures, and the team confident enough to change prompts without anxiety.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
Collect twenty real cases the system got wrong and write down what the right answer would have been. That is your first evaluation set, and it is more useful than a hundred cases someone invented.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: evaluation sets built from real inputs and extended from production failures, rubrics tested for agreement before they are relied on, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read AI evaluation cost.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What should the first evaluation set contain?
Real inputs the system will see, weighted towards the difficult cases and the failures you already know about. A set of easy cases the system handles well measures nothing and produces false confidence.
02How do you know the rubric works?
By having two people rate the same outputs independently and comparing. Low agreement means the rubric is ambiguous rather than the raters being careless, and the fix is clarifying the rubric with worked examples.
03How often should evaluation run?
On every change to prompts, models, retrieval corpora, or configuration, automatically. Quarterly evaluation tells you the system regressed at some point in the last three months, which is not actionable.
04Why do evaluation programmes fail?
Because nobody acts on the results. A suite that reports a regression into a dashboard nobody watches, with no owner and no gate, is documentation of decline rather than a control.
05How should the set grow?
From production failures. Every case the system got wrong in production becomes a test case, which means the suite becomes progressively better at catching the things that actually break rather than the things someone imagined.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.