FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Decision Guide · 4 minute read

How to Write Acceptance Criteria for AI Systems

Acceptance criteria for AI are measurable statements of what the system must achieve to be accepted: accuracy or task success thresholds by category on a named golden dataset, groundedness and safety test results, latency and cost limits at defined load, behavior on low confidence and failures, operability requirements, and documentation deliverables, each with an evaluation method and an owner.

By FISTA Solutions· AI-Native Engineering Team·
How to Write Acceptance Criteria for AI Systems article cover

Ask a team when an AI system is done and most will say when it works. Ask what works means and the answers diverge: for the engineer, the demo runs; for the user, it never embarrasses them; for the sponsor, the number moves. Acceptance criteria resolve that by stating, before build, what the system must measurably achieve, on which dataset, judged how, signed off by whom. This guide covers how to write them with examples, drawing on FISTA Solutions' AI enablement practice. The specification they belong to is in how to write an ai spec and the testing process in the ai acceptance testing checklist.

What makes a criterion acceptable?

PropertyWeak criterionStrong criterion
MeasurableAccurate answersTask success at or above threshold on the golden dataset
ScopedOverallBy category, with thresholds per category
EvidencedTeam demoEvaluation report with method stated
OwnedThe projectNamed sign-off per criterion
BoundedFast enoughTime to first token and completion at defined load
CompleteQuality onlyQuality, safety, cost, failure behavior, operability, documentation

Which categories of criteria should a specification include?

  • Quality: task success or accuracy thresholds by category on the named golden dataset, with the grader and calibration stated. Dataset design is in what is a golden dataset.
  • Groundedness: for retrieval systems, the share of claims supported by cited sources. See what is groundedness in ai.
  • Safety: adversarial and policy test suites passed at defined rates. See what is ai red teaming.
  • Failure behavior: what the system does on low confidence, missing data, and tool errors, verified by test cases.
  • Performance: latency and throughput at defined load. See what is latency in ai systems.
  • Cost: cost per task at expected volume within a ceiling.
  • Integration: actions correct, permissions enforced, idempotency verified.
  • Operability: tracing, monitoring, alerts, runbooks, rollback tested.
  • Documentation: model or system card, evaluation report, handover. See what is a model card.

How do you set thresholds?

Per category by consequence: higher where errors are costly, irreversible, or customer-visible; lower where a human reviews every output anyway. Anchor them in the baseline the system replaces, in what discovery evaluation showed reachable, and in the autonomy level planned. Record the reasoning so thresholds can be revisited when evidence changes. One aggregate accuracy number hides the category that fails. Evaluation practice is in the ai evaluation checklist.

What do worked examples look like?

  • Support triage agent: routing accuracy at or above threshold per ticket category on the golden dataset of labeled tickets, graded by exact match to the taxonomy; low-confidence tickets routed to human triage at or below a defined rate; median time to route under a limit at peak load; cost per ticket under a ceiling; runbook and rollback tested; business owner signs off on accuracy, engineering on operability.
  • Document extraction pipeline: field-level accuracy thresholds per document type and field on the labeled set; fields below confidence routed to review; throughput at defined volume; per-document cost ceiling; validation rules verified on adversarial documents.
  • Knowledge assistant: groundedness share on the question set; refusal on out-of-scope questions verified; permission leakage tests passed at zero; time to first token under limit; model card delivered.

Domain builds are in how to build an ai ticket routing system and how to build an ai data extraction pipeline.

Who signs off, and against what?

A named business owner for quality and failure behavior; evaluation or engineering for method, performance, and operability; security for safety criteria; governance where the risk tier requires review. Sign-off is against evidence presented per criterion, in a report, never against a demo. Sign-off records become the acceptance record in the contract. Contract linkage is in what is a statement of work.

How do criteria live on after acceptance?

They become the release gate: every prompt, model, retrieval, or tool change re-runs the evaluation against the same criteria before production. Production sampling checks the criteria continue to hold. Changes to the criteria themselves are recorded decisions with reasons. Gate implementation is in how to build an ai quality gate.

What mistakes weaken criteria?

Aggregate thresholds; criteria without datasets; safety and cost omitted; failure behavior unspecified; no named sign-off; criteria written after the build to match what it does; and criteria that are never re-run after acceptance. Each turns acceptance back into an argument.

How FISTA Solutions writes acceptance criteria

FISTA Solutions writes acceptance criteria during discovery with client domain experts, ties each to a golden dataset and evaluation method, sets thresholds per category by consequence, names sign-off owners, and wires the criteria into CI as the release gate for every later change. The AI enablement practice leads specification and evaluation, forward deployed engineers deliver against the criteria, and AI agents ship gated by them. The record behind the approach is 150+ projects with 99.9% uptime.

To define done before you build, message FISTA on WhatsApp, or read how to write an ai spec for the document acceptance criteria belong to.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01Why can't AI acceptance be works as expected?

Because outputs vary and expectations differ by person. A probabilistic system is accepted against measured thresholds on an agreed dataset, so that everyone sees the same evidence and the same criteria gate every future change. Impressions produce disputes; measurements produce decisions.

02How do you set thresholds?

Per category by consequence: higher thresholds where errors are costly or irreversible, lower where a human reviews anyway. Ground them in the baseline the system replaces and in what discovery evaluation showed reachable. Record the reasoning so thresholds can be revisited.

03What categories of criteria should exist?

Quality by task category, groundedness for retrieval systems, safety and adversarial test results, latency and throughput at defined load, cost per task, behavior on low confidence and errors, integration correctness, operability with monitoring and runbooks, and documentation including model cards.

04Who signs off?

A named business owner for quality and behavior criteria, evaluation or engineering for method and operability, security for safety criteria, and governance where risk tier requires it. Sign-off is against evidence presented per criterion, not against a demo.

05How do criteria evolve?

They start in discovery from what correct means, are refined as the golden dataset grows, become the pilot's success definition, gate production release, and remain the regression gate for every prompt, model, or data change afterward. Changes to criteria are recorded decisions.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project