Playbook · 6 minute read
How to Monitor AI Quality in Production Continuously
Monitoring AI quality in production combines proxy signals that correlate with quality, sampled human review, drift detection on the input distribution, and evaluation runs on a schedule. Availability monitoring will report that a system producing confidently wrong answers is perfectly healthy.
Availability monitoring reports that a system producing confidently wrong answers is healthy. Monitoring AI quality needs different signals entirely. This playbook covers what to measure and what to alert on, drawing on FISTA Solutions' AI enablement work.
When is this worth doing?
From the day an AI system reaches production. Retrofitting monitoring after a quality incident means the incident was found by users, which is the expensive route to the same conclusion.
It is also worth revisiting when a system's audience or input distribution changes materially, because the proxies and thresholds may no longer be calibrated.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Instrument proxy signals | Free, continuous, correlated |
| 2. Set up sampled review | The only direct measure |
| 3. Detect input drift | It predicts decline |
| 4. Schedule evaluation runs | Not only on change |
| 5. Alert on rate changes | Not absolute values |
| 6. Tie alerts to action | Or they get ignored |
Step 1 â Instrument the proxy signals
Escalation rate, user rephrasing, session abandonment, heavy editing of outputs, retry rate, abstention rate, and tool call patterns.
These cost nothing per interaction and give continuous coverage. A rise in rephrasing means users are not getting what they asked for; a rise in editing means the output needed work before it was usable.
Establish the baseline for each while the system is known to be performing acceptably. Without a baseline, a rate is a number rather than a signal.
Step 2 â Set up sampled human review
A small daily sample: random cases plus every escalation, every abstention, and every unusually expensive trajectory.
This is the only direct measure of quality. Proxies correlate imperfectly, and a system can be wrong in ways users accept â a plausible summary with an error nobody checked â which no proxy detects.
Make the review interface fast. Sampling that takes an hour a day stops happening within a fortnight; sampling that takes fifteen minutes continues indefinitely.
Step 3 â Detect input drift
Monitor the distribution of inputs: topics, lengths, languages, formats, and request types.
Drift predicts quality decline. A system evaluated on one distribution and serving another will perform differently even if nothing inside it changed, and the change frequently arrives gradually enough that nobody notices.
Simple measures work: the proportion of requests falling into each known category, and the proportion falling outside all of them. A rising unknown category is an early warning. See what is continuous evaluation.
Step 4 â Schedule evaluation runs
Run the evaluation suite on a schedule as well as on every change.
The corpus can change, the provider's model can change, and the input distribution can shift, all without a deployment. Change-triggered evaluation alone misses every one of those.
Weekly is usually enough for a stable system, daily for one under active change or in a high-consequence domain.
Step 5 â Alert on rate changes
At production volume, individual failures are constant and meaningless. The signal is a change in rate.
Alert on escalation rate rising, abstention rate changing in either direction, evaluation score dropping below threshold, cost per task rising, and unknown-category inputs increasing.
Set thresholds from the baseline rather than from intuition, and review them as the system matures. Thresholds set at launch are frequently wrong within a quarter.
Step 6 â Tie every alert to an action
Each alert should have a runbook entry: what it means, what to check, what to do.
Alerts without actions get muted. Within a month of an alert firing repeatedly with no clear response, the team stops looking at it, and the alert that mattered is in the same channel as the noise.
Review alert usefulness monthly: what fired, what needed action, what was noise. Remove or tune the noise. See how to build an ai runbook.
Can models monitor quality automatically?
Partly, and with validation. Model-based scoring works for some criteria â format compliance, presence of required elements, obvious contradictions with retrieved context â and poorly for domain judgement.
Validate any automated scorer against human ratings before relying on it, and re-validate when models change. A scorer that drifts produces a quality signal that is itself wrong, which is worse than no signal because it looks like data.
What about monitoring for unsafe output?
Separately and with a lower tolerance. Unsafe or inappropriate output warrants immediate alerting regardless of rate, because a single instance can be a serious event.
That needs its own detection: classifiers, keyword screens, or policy checks at the output boundary, tuned to catch more rather than fewer. False positives here are cheap; misses are not.
Who needs to be involved?
An engineer to instrument it, someone with domain knowledge for the sampled review, and a named owner for the quality number.
The domain reviewer is the constraint and the part most often unassigned, which is why sampled review is the first thing to lapse.
How long does it take?
One to two weeks to instrument proxies and set up sampling, then continuous. Baselines need a few weeks of normal operation before they mean anything.
What are the common failure modes?
Monitoring availability only. Proxies without baselines. Sampled review that takes too long. No drift detection. Alerting on absolute values. And alerts with no runbook entry.
How do you know it worked?
Quality decline detected by the team before users report it, alert volume low and actionable, sampled review still happening after three months, and evaluation catching provider-side changes.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
Instrument escalation rate and output editing rate, and establish their baselines. Those two proxies catch a surprising proportion of quality problems for very little effort.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: proxy signals instrumented with baselines alongside sampled human review, drift detection on inputs so decline is predicted rather than discovered, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read how to run AI operations daily.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What are proxy signals?
Behaviour that correlates with quality without needing a human judgement: escalation rate, user rephrasing, abandonment, heavy editing of outputs, retry rate, and abstention rate. They cost nothing and run continuously.
02Why is sampled review still needed?
Because proxies correlate imperfectly. A system can be wrong in ways users do not notice and therefore do not escalate, and only someone looking at outputs with domain knowledge catches that class of failure.
03What is input drift?
A change in the distribution of what users are asking or submitting. It predicts quality decline, because a system evaluated on one distribution and serving another will perform differently regardless of whether anything inside it changed.
04What should alerts fire on?
Rate changes rather than absolute values. At production volume individual failures are constant; a shift in the escalation rate or the abstention rate is the signal that something has actually changed.
05How often should the evaluation suite run in production?
On a schedule as well as on change, because the corpus and the provider's model can shift without any deployment on your side. Weekly is usually enough to catch drift before users report it.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.