FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Playbook · 6 minute read

How to Set AI Quality Thresholds People Will Act On

Quality thresholds work when they are anchored to the process being replaced rather than to an arbitrary score, set per category rather than in aggregate, and attached to a consequence — a blocked release or an escalation — rather than reported into a dashboard.

By FISTA Solutions· AI-Native Engineering Team·
How to Set AI Quality Thresholds People Will Act On article cover

Quality thresholds only matter if something happens when they are breached. This playbook covers setting thresholds anchored to reality, differentiated by consequence, and attached to a control, drawing on FISTA Solutions' AI enablement work.

When is this worth doing?

Before a system reaches production, and whenever the input distribution or the use case changes materially.

It is also worth doing for systems already in production without thresholds, which is common. Those systems are being judged by whoever complains, which is not a standard anyone can build to.

What does the sequence look like?

StepPurpose
1. Measure the current processThe honest comparison
2. Segment by consequenceDifferent bars for different stakes
3. Set thresholds per categoryNot in aggregate
4. Define below-threshold behaviourAbstain or escalate
5. Attach consequencesGates, not notifications
6. Revisit as things changeDistribution, use, and stakes

Step 1 — Measure the current process

Find out how well the existing process performs: error rate, consistency between people, and how often mistakes are caught downstream.

This is the honest comparison and it is usually more forgiving than teams expect. Human performance on repetitive judgement tasks is rarely as good as the assumption behind a ninety-nine per cent target.

It also grounds the conversation. A threshold justified by measured current performance is defensible; one justified by a round number invites an argument nobody can settle.

Step 2 — Segment by consequence

Group request types by what happens when the system is wrong: inconvenience, rework, financial loss, regulatory exposure, or harm.

Those groups deserve different bars. Applying one threshold across all of them is either too permissive for the serious cases or too strict for the trivial ones, and usually both.

This segmentation also tells you where human review is warranted, which is frequently a better control than a higher threshold.

Step 3 — Set thresholds per category

Give each category its own threshold, derived from the current process and the consequence.

Aggregate thresholds hide category failures. A system at ninety-two per cent overall may be at ninety-nine on the common cases and sixty on an important minority, and the average conceals exactly the problem that matters.

State the thresholds in terms someone outside the team can understand: how often the system is wrong, on which kinds of case, and what happens then.

Step 4 — Define what happens below threshold

A system that cannot meet the bar for a case should abstain or escalate rather than answer.

That requires the system to know when it is uncertain, which means confidence signals, retrieval coverage checks, or explicit refusal conditions. Building those is part of meeting the threshold rather than separate from it.

Abstention is underused. A system that declines five per cent of cases and is right on the rest is more useful than one that answers everything and is wrong on five per cent, because the first failure mode is visible. See what is abstention in ai.

Step 5 — Attach consequences to breaches

Decide what happens when a threshold is breached: the release is blocked, traffic routes to a fallback, the category escalates to humans, or the system is disabled.

A threshold without a consequence is a report, and reports get muted. The discipline of deciding the consequence also forces a realistic threshold, because nobody sets a bar that would block every release.

Automate the consequence where possible. Thresholds enforced by a person noticing depend on that person noticing.

Step 6 — Revisit as things change

Input distributions widen, use cases expand, and consequences change as a system becomes more embedded.

A threshold set at launch for a narrow audience may be wrong once the system serves everyone. Review it when the audience changes, when a new category of request appears, and when the system starts being relied on more heavily than it was.

Record the reasoning each time. Thresholds that drift without a record become numbers nobody can justify.

What about thresholds nobody can meet?

They are a signal that the use case is wrong for the technology, or that the scope is too broad.

A system that cannot reach an acceptable bar across a wide range of cases may reach it easily on a narrower one. Narrowing scope is usually a better response than accepting a bar the business does not actually find acceptable.

The worst outcome is quietly lowering the threshold to match what the system achieves, which converts a quality standard into a description.

How do thresholds relate to error budgets?

Closely. A threshold says what quality is acceptable; an error budget says how much of the allowance has been consumed and what happens when it runs out.

That framing works well for AI systems because it makes the trade explicit: a team with budget remaining can ship faster, and one that has exhausted it stabilises. See what is an error budget.

Who needs to be involved?

Someone who owns the business outcome, someone who can measure quality, and whoever will enforce the consequence.

Thresholds set by engineering alone tend to be technically defensible and commercially wrong. Thresholds set by business alone tend to be unmeasurable.

How long does it take?

One to two weeks, mostly measuring the current process. The threshold decisions themselves take an afternoon once the measurement exists.

What are the common failure modes?

Round numbers with no basis. Aggregate thresholds. No below-threshold behaviour. Breaches that only notify. And quietly lowering the bar to match performance.

How do you know it worked?

Thresholds someone outside the team can explain, breaches producing action rather than notification, and abstention behaviour that users find reasonable.

What does it cost?

Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.

Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.

What should you do first?

Measure how often the current human process gets this task wrong. That number anchors every threshold conversation that follows.

How FISTA Solutions helps

FISTA Solutions runs this work alongside client teams rather than around them: thresholds anchored to measured current performance and differentiated by consequence, breaches wired to gates rather than dashboards, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.

To run this with support, message FISTA on WhatsApp, or read what is an error budget.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01What should thresholds be anchored to?

The process being replaced or assisted. If humans doing this task are right ninety-four per cent of the time, that is the comparison. Arbitrary targets like ninety-five per cent produce arguments rather than decisions.

02Why per category?

Because an aggregate score hides category failures. A system scoring well overall while failing an entire class of important cases has regressed in a way the average cannot show, and users in that class experience it as broken.

03What makes a threshold real?

A consequence. A breach that blocks a release, triggers an escalation, or routes traffic away is a control. A breach that appears on a dashboard is information nobody is obliged to act on.

04Should all cases have the same bar?

No. Consequence should set the bar. A wrong answer about opening hours is an inconvenience; a wrong answer about a medication or an entitlement is not, and the same threshold for both is wrong in one direction or the other.

05What happens below the threshold?

The system should abstain or escalate rather than answer. A confident wrong answer is worse than no answer, and designing the below-threshold behaviour is as important as setting the threshold.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project