FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Glossary ┬╖ 5 minute read

What Is Continuous Evaluation? Ongoing AI Quality Measurement

Continuous evaluation measures AI quality on an ongoing basis in production rather than once before launch. It combines automatically computable signals, sampled human review, and periodic refresh of the evaluation set from real traffic, because models, data, and usage all change after deployment.

By FISTA Solutions┬╖ AI-Native Engineering Team┬╖
What Is Continuous Evaluation? Ongoing AI Quality Measurement article cover

Most AI systems are evaluated once, before launch, against a set assembled at the time. Everything then changes тАФ the model, the content, the users, the questions тАФ while the system continues to be judged against that frozen snapshot, if it is judged at all. This explainer covers what ongoing measurement looks like. It complements what is offline vs online evaluation and ai evaluation vs ai monitoring, and reflects FISTA Solutions' approach in AI enablement delivery.

Why does one-off evaluation decay?

Because every input to quality moves. The provider updates the model. The document corpus gains and loses content. A product launch changes what users ask. A marketing campaign brings a different population. A prompt is tweaked.

Each of these changes behaviour without triggering any alert, and a system that scored well six months ago is being described by a measurement of a configuration and a user base that no longer coexist.

SignalAutomatableFrequency
Citation validityYesEvery request
Schema conformanceYesEvery request
Abstention and refusal rateYesEvery request
Retrieval recallYes, against labelled setScheduled
Task completionUsuallyEvery request
Helpfulness and toneNoSampled review

What can be measured automatically?

More than most teams instrument. Whether every factual claim carries a citation and whether the cited source supports it. Whether output conforms to the required schema. Abstention and refusal rates. Retrieval recall against a labelled query set. Task completion. Latency and cost. And implicit behavioural signals: retries, rephrasing, abandonment, escalation.

These run on every request at negligible cost and catch most degradation before a user complains.

Where is human review necessary?

For judgements automation cannot make. Whether an answer was genuinely helpful for what the person needed. Whether the tone was right. Whether a subtle factual error is present that no automated check would catch.

Sampling makes this affordable, and the sampling should be stratified rather than uniform: oversample low-confidence responses, escalated cases, and new query types, because that is where problems concentrate.

How is the evaluation set kept current?

By sampling production traffic, labelling it, and adding it on a schedule. That loop is what keeps offline evaluation predictive, and it is the step most often omitted because it costs ongoing human effort.

Budgeting that effort explicitly is the difference between a set that stays useful and one that becomes a ritual gate nobody trusts. See what is offline vs online evaluation.

What causes drift?

Provider model updates, which can change behaviour without notice. Content changes in the corpus. Shifts in what users ask. Prompt edits. Dependency behaviour changes, such as a retrieval service returning differently ranked results.

None of these affects availability, which is precisely why conventional monitoring reports a healthy system throughout.

What should be alerted on?

Rate-of-change on the automatable signals, not absolute thresholds alone. A citation validity rate that drops four points in a day is a signal even if it remains above target, because something changed and the change was not intentional.

What should you do first?

Pick one automatable signal тАФ citation validity or schema conformance тАФ and measure it continuously on production traffic for a month. The variation you observe usually settles the argument about whether continuous evaluation is worth building.

How does it fit with incident response?

Directly. A quality signal that moves sharply is an incident, and treating it as one тАФ with an owner, an investigation, and a resolution тАФ is what distinguishes monitored systems from measured ones. Most organisations have incident processes for availability and nothing equivalent for quality.

What does it cost?

The automated signals are cheap. The recurring cost is human review of the sample and the labelling that keeps the evaluation set current, which is a standing commitment of expert time. Budgeting it as a running cost rather than a project expense is what keeps it from being cut after the first quarter.

Who reviews the sample?

People with domain knowledge, not a general reviewing pool. Judging whether a clinical summary or a policy answer was genuinely correct requires knowing the domain, and a reviewer without it can only assess fluency, which is the quality least in need of checking.

Scheduling that expert time is the practical constraint, and it is what determines how large a sample is realistic.

What does it look like operationally?

A dashboard with the automated quality signals alongside availability and latency, alerts on rate of change, a weekly sample review with domain experts, and a scheduled refresh of the evaluation set from production traffic.

None of that is elaborate. What makes it work is that someone owns it and looks at it, which is the difference between measurement and instrumentation.

How FISTA Solutions helps

FISTA Solutions instruments automatable quality signals on every production request, stratifies human review sampling toward where problems concentrate, refreshes evaluation sets from real traffic on a schedule, alerts on rate of change rather than thresholds alone, and budgets ongoing labelling as a running cost, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.

To know your AI quality today rather than at launch, message FISTA on WhatsApp, or read ai evaluation vs ai monitoring.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01Why does one-off evaluation decay?

Because everything moves. The provider updates the model, the document corpus changes, users ask different things, and a new product introduces new question types. A score from six months ago describes a system and a population that no longer exist together.

02What can be measured automatically?

Grounding and citation validity, schema conformance, abstention and refusal rates, retrieval recall against known answers, task completion, latency, cost, and implicit signals such as retries and abandonment. These run on every request at negligible cost.

03Where is human review needed?

For judgements automation cannot make: whether an answer was genuinely helpful, whether tone was appropriate, whether a subtle factual error is present. Sampling makes this affordable, and it should be stratified rather than uniform.

04How is the evaluation set kept current?

By sampling production traffic, labelling it, and adding it to the set on a schedule. Without that, the set describes an increasingly historical picture of usage while continuing to gate every change.

05What causes drift?

Provider model updates, changes to the underlying content, shifts in what users ask, prompt changes, and dependency behaviour changes. Each affects quality without affecting availability, which is why conventional monitoring misses all of them.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.

Start a project