Leadership · 4 minute read
How to Read an AI Evaluation Report
Read an AI evaluation report by checking the set before the score: how many cases, where they came from, who labeled them, and whether they represent production inputs. Then read the failure analysis, which is more informative than the pass rate, and check the date and the threshold.
An evaluation report is the main evidence executives receive about whether an AI system works, and most are read badly: the eye goes to the headline percentage, which is the least informative number on the page. This guide covers how to read one properly, in the order that matters.
Read the set before the score
The pass rate is a property of the test set as much as of the system. Five questions determine what it means:
| Question | Strong answer | Weak answer |
|---|---|---|
| How many cases? | Hundreds or more for a production system | A dozen |
| Where did they come from? | Sampled from real production inputs | Invented by the build team |
| Who labeled the correct outcomes? | People who own the process | The engineers, or the model itself unchecked |
| Are hard and adversarial cases included? | Yes, deliberately | Only clean examples |
| How has it grown? | Every production failure added | Static since it was created |
A 95% pass rate on twenty invented clean cases is worse evidence than 85% on five hundred real ones including the awkward tail. The AI evaluation explained for executives piece covers what good evaluation practice looks like.
Read the failure analysis next
More informative than the headline. The questions: what kinds of case failed, how severe were the failures, and were they concentrated or scattered?
- Failures on rare, low-consequence cases: usually acceptable, and they tell you where the escalation rules need to be.
- Failures on high-consequence cases: the pass rate is irrelevant; those cases need human review regardless.
- Scattered unpredictable failures: worrying, because they cannot be bounded by a rule.
- Failures clustered on one input type: fixable, and probably a context or data problem.
Two systems with the same pass rate can have entirely different risk profiles, and only the failure analysis distinguishes them. The AI agent failure modes for executives piece covers the categories.
Check the date and cadence
AI behavior changes when the model, the prompts, the retrieved content, or the input distribution changes, none of which requires a code deployment. An evaluation run three months ago describes a system that may no longer exist.
Ask when it was last run, on what cadence it runs in production, and whether anything has changed since: model version, prompts, tools, or upstream data sources. A report presented without a date is not evidence.
Check the threshold and its timing
Two questions: what threshold was set, and when. A threshold defined before the result was known is a standard; one defined afterward is a rationalization. Ask whether anything has shipped below threshold and what justified it.
The threshold itself should come from the business owner based on consequence, not from engineering based on what was achievable. The how much autonomy should AI agents have guide covers the relationship between evaluation evidence and the authority granted.
Check what was judged and by whom
"Correct" is a judgment. Ask what criteria were used, who applied them, and whether a model was used to judge model output, which is common and acceptable when validated against human judgment but misleading when not. For subjective outputs, ask about agreement between human judges; low agreement means the criteria are unclear and the score is soft.
What is not in most reports and should be?
- Production comparison: how the live system performs on sampled real traffic, versus the test set.
- Trend: the same set, run over time, showing direction.
- Coverage statement: what was not tested.
- Adversarial results: how it behaved under deliberate attempts to break it.
Requesting these once changes what teams produce thereafter, which is one of the more efficient interventions available to an executive. The AI red teaming explained for executives piece covers the adversarial half.
How should reports change as a system matures?
Early reports are about readiness: is this good enough to deploy under supervision? Later reports are about stability: has anything moved, and why? The metrics stay the same and the emphasis shifts to the trend line and to the comparison between test-set performance and sampled production behavior, which is where drift appears first. A mature system's report should be dull, and a report that suddenly becomes interesting is telling you something changed.
What should executives ask?
- How many cases, from where, labeled by whom?
- What failed, and how bad were those failures?
- When was this run, and what has changed since?
- What threshold applied, and was it set before or after?
- How does the live system compare with the test result?
How can FISTA Solutions help?
FISTA Solutions builds evaluation sets from clients' real production cases, labeled with the people who own the process, with failure analysis, adversarial cases, and scheduled re-runs, as part of every AI agent engagement, and its Applied division reviews evaluation evidence for systems built elsewhere. Since 2017, FISTA has delivered 150+ projects for 50+ companies across 12+ countries.
To have an evaluation report independently reviewed before you rely on it, talk to FISTA on WhatsApp, or read how executives should evaluate an AI demo.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What does an AI pass rate actually mean?
The share of cases in a specific test set where the system produced an outcome judged correct, under the judging criteria used. It is only as meaningful as the set is representative and the criteria are sound, which is why the set description matters more than the number itself.
02How do you tell if an evaluation set is any good?
Ask how many cases it contains, where they came from (real production inputs or invented examples), who labeled the correct answers and with what expertise, whether hard and adversarial cases are included, and how it has grown from production failures. A small set of easy invented cases proves nothing.
03Why does the failure analysis matter more than the pass rate?
Because the kind of failure determines the risk. A system failing on rare, low-consequence cases at the same rate as one failing unpredictably on high-consequence cases has an identical pass rate and a completely different risk profile. Read what failed and how badly.
04What threshold should an AI system meet?
One set in advance by the business owner, based on the consequence of errors and what catches them. There is no universal number. The important question is whether the threshold was defined before the result was known, and whether anything shipped below it.
05How current does an evaluation need to be?
Recent, and repeated on a schedule. AI behavior changes when models, prompts, data, or inputs change, so a result from three months ago describes a system that may no longer exist. Check the run date and the cadence of scheduled re-evaluation.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.