FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Playbook · 6 minute read

How to Write AI Acceptance Criteria That Can Be Tested

AI acceptance criteria work when they specify behaviour on defined categories of input with measurable thresholds, state explicitly what the system must refuse or escalate, and include operational requirements such as evaluation infrastructure, logging, and a working escalation path rather than output accuracy alone.

By FISTA Solutions· AI-Native Engineering Team·
How to Write AI Acceptance Criteria That Can Be Tested article cover

Acceptance criteria written for deterministic software do not work for AI, because the system's behaviour is statistical rather than specified. This playbook covers writing criteria both sides can test, drawing on FISTA Solutions' forward deployed engineers work.

When is this worth doing?

Before any AI project with a defined deliverable — internal or with a supplier — and particularly where payment or sign-off depends on acceptance.

It is also worth doing for internal work with no contract, because the discipline of writing testable criteria surfaces disagreements about what the system is for while they are still cheap.

What does the sequence look like?

StepPurpose
1. Define input categoriesWhat the system will see
2. Specify behaviour per categoryIncluding refusals
3. Set measurable thresholdsPer category, with a basis
4. Agree the evaluation setBefore building starts
5. Add operational criteriaLogging, monitoring, escalation
6. Define the acceptance testWho runs it and when

Step 1 — Define the input categories

List the kinds of input the system will receive, with real examples of each.

That list is the foundation. Criteria written without it default to overall accuracy, which is both untestable and uninformative — a system can hit ninety per cent overall while failing an entire category.

Include the categories the system should handle badly or not at all: out-of-scope requests, adversarial inputs, and things requiring human judgement. Those need criteria too.

Step 2 — Specify behaviour per category

For each category, say what the system should do: answer, answer with a caveat, escalate, or refuse.

This is more useful than accuracy because it is unambiguous. A criterion stating that requests involving pricing must escalate rather than answer is testable by anyone; a criterion about accuracy on pricing questions is not.

Refusal criteria are the ones most often missing and the ones that prevent the worst outcomes. A system that must refuse to give legal or medical advice needs that written down and tested.

Step 3 — Set measurable thresholds with a basis

Give each category a threshold derived from the process being replaced or from the consequence of error.

Round numbers with no basis produce arguments at acceptance. A threshold justified by measured current performance is defensible to both sides.

State the measurement method too: how many cases, drawn from where, scored by whom against what rubric. Ambiguity there is where acceptance disputes actually happen. See how to set ai quality thresholds.

Step 4 — Agree the evaluation set before building

Both parties should agree the cases the system will be measured against, before development starts.

This is the single most effective step. It forces the conversation about what the system is for, it gives the builder a target, and it removes the possibility of the test being chosen after the fact by whoever is unhappy.

Hold back a portion the builder does not see, to avoid the system being tuned to the test. Agree that split up front so it does not look like a trick later.

Step 5 — Add operational criteria

Evaluation infrastructure that can be re-run, logging sufficient to reconstruct any decision, monitoring that detects quality drift, a working escalation path, and documentation.

Those are what separate a system you can operate from a demonstration that passed a test. Criteria covering only output quality produce systems that work on day one and cannot be maintained.

Include handover explicitly: what documentation, what runbooks, and a session where the receiving team changes something themselves.

Step 6 — Define the acceptance test itself

Say who runs it, when, on what environment, with what data, and what happens if it fails.

Acceptance processes that are undefined become negotiations. A written process — the evaluation suite runs on the held-back set, in this environment, scored by these people, with these thresholds — is settled in advance.

Define partial acceptance too. Most systems pass some categories and not others, and having agreed what that means avoids an all-or-nothing argument.

How do you handle behaviour you cannot anticipate?

With a category for unanticipated inputs and a behavioural criterion: the system should abstain or escalate rather than guess.

That is testable — feed it inputs outside its scope and check it declines. It is also the right behaviour, and specifying it prevents the common outcome where a system confidently handles something nobody intended it to see. See what is abstention in ai.

What about criteria that change during the project?

Expect some. AI projects surface things about the data and the task that nobody knew at the start.

Handle it with a change process rather than by ignoring the criteria: revise them explicitly, with both sides agreeing, and record why. Projects where criteria quietly drift end in disputes about what was agreed.

A discovery that a category is harder than expected is legitimate grounds for revision. A discovery that the system cannot meet the threshold is not.

Who needs to be involved?

Someone who knows what correct looks like, someone who will operate the system, and the builder.

Criteria written by the builder alone describe what they intend to build. Criteria written by the buyer alone frequently describe something no system can do.

How long does it take?

One to two weeks, mostly spent assembling the evaluation set and agreeing categories. That time is recovered several times over during delivery.

What are the common failure modes?

Overall accuracy criteria. No refusal behaviour. Thresholds with no basis. Evaluation set agreed after building. No operational criteria. And an undefined acceptance process.

How do you know it worked?

An acceptance test both parties can run and agree on, no dispute about what was promised, and a system that is operable rather than merely accurate.

What does it cost?

Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.

Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.

What should you do first?

Write down five categories of input the system will see and what it should do with each. That list is the skeleton of everything else.

How FISTA Solutions helps

FISTA Solutions runs this work alongside client teams rather than around them: acceptance criteria written per input category with agreed evaluation sets, operational requirements included alongside output quality, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.

To run this with support, message FISTA on WhatsApp, or read how to scope an AI agent project.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01Why do conventional acceptance criteria fail for AI?

Because they assume deterministic behaviour. A criterion stating the system shall correctly categorise incoming requests is untestable without saying which requests, what correct means, and at what rate.

02What should criteria specify?

Input categories, the expected behaviour for each, a measurable threshold, and the evaluation set the measurement uses. That combination is testable by either party without argument at acceptance.

03Why include refusal behaviour?

Because a system that answers everything will answer things it should not. Criteria specifying what must be refused or escalated are as important as those specifying what must be answered correctly.

04What operational criteria belong?

Evaluation infrastructure, logging sufficient to reconstruct a decision, monitoring for quality drift, a working escalation path, and documentation. Those are the difference between a demonstration and something operable.

05What about criteria that cannot be measured?

Either make them measurable or remove them. A criterion requiring the system to be helpful and professional cannot be tested and becomes an argument at acceptance, which serves neither party well.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project