FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Cost · 5 minute read

AI Evaluation Cost: Building Sets, Grading and Keeping Current

AI evaluation cost is dominated by human judgement — constructing labelled sets and grading outputs — rather than by compute. Model judges reduce grading cost where criteria are checkable, and every evaluation set decays as usage changes, which makes refresh a recurring rather than one-off cost.

By FISTA Solutions· AI-Native Engineering Team·
AI Evaluation Cost: Building Sets, Grading and Keeping Current article cover

Evaluation budgets are frequently built around compute, which is the smallest component. The cost is human: deciding what good looks like, constructing a set that represents real usage, and grading outputs against criteria consistently. This guide covers what actually drives the number, drawing on FISTA Solutions' AI enablement work. It complements what is an evaluation rubric and what is continuous evaluation.

Why does human judgement dominate?

Because someone must decide what good means and whether each output achieves it. Running a model against a prepared set is inexpensive; preparing the set and grading the outputs is expert time.

That inverts the intuitive cost model, and it explains why evaluation is consistently under-resourced: the visible cost is small and the real cost is a demand on people who are already busy.

ComponentCostFrequency
Criteria definitionModerateOne-off, revised
Set constructionLargeOne-off per task
Human gradingLargeRecurring
Model-judge validationModeratePer task
Set refresh from productionModerateRecurring
Compute to run evaluationsSmallEvery change

What makes set construction expensive?

Selecting cases that represent actual usage, determining the correct answer or acceptable range for each, and deliberately covering failure modes — including the questions the system should decline.

That last category takes the most thought and is omitted most often. A set containing only answerable questions cannot detect a system that has started answering things it should not, which is a serious regression that stays invisible.

When do model judges reduce cost safely?

On checkable criteria. Is a citation present, does the cited source support the claim, does the output conform to the schema, is a required element included. These have determinate answers and a model judge scales to volumes human grading cannot reach.

On subjective judgements — was this genuinely helpful, was the tone right — judges carry biases toward length and toward their own style. They can still be used, after measuring agreement against human grades on that specific task, and a judge validated for one task does not transfer to another.

Why do evaluation sets decay?

Because traffic changes. New products introduce question types nobody anticipated, a campaign brings a different user population, and a feature launch shifts what people ask about.

A set assembled six months ago describes six-month-old usage while continuing to gate every change. Refreshing it from sampled production traffic is what keeps it predictive, and it is recurring labelling work that should be budgeted as such.

How does domain expertise affect cost?

Sharply. Grading a clinical summary, a legal analysis, or an engineering assessment requires someone with the domain knowledge, and that time is scarce as well as expensive.

Task design matters here as it does in labelling: separating the checkable properties from the judgement-requiring ones lets automation and non-experts handle the first, reserving expert time for what genuinely needs it.

What can reduce the cost?

Stratified sampling rather than uniform grading, so expert attention goes to low-confidence and high-consequence outputs. Automating the checkable criteria. Reusing sets across related tasks where the criteria overlap. And building the set incrementally from real failures rather than attempting comprehensive coverage up front.

That last approach produces a set that reflects how the system actually fails, which is more useful than one reflecting how someone imagined it might.

Why is no evaluation the most expensive option?

Because it means shipping blind. Regressions reach users, are discovered through complaints, and are diagnosed without a baseline to compare against — which turns a measurement problem into an investigation.

That cost is diffuse and therefore consistently under-weighted in budget discussions, while the evaluation cost is concrete and visible. The comparison is unfair in a predictable direction.

What should you do first?

Take your last user-reported failure and turn it into an evaluation case. That is a set with one entry, it costs almost nothing, and it establishes the practice that grows the set in the direction the system actually fails.

How does evaluation cost scale with the estate?

Sub-linearly if the harness is shared, linearly if it is not. Organisations that build evaluation infrastructure once — case management, grading workflow, result storage, comparison reporting — add each new system's evaluation at the marginal cost of its set and its criteria.

Organisations that build evaluation per project rebuild the harness every time and pay the full cost repeatedly. That difference compounds quickly across an estate and is the strongest argument for treating evaluation as platform capability rather than as project work.

Who should do the grading?

People with the domain knowledge for the judgement being made, which for many enterprise systems means the people who currently do the work the system is assisting with. That has a secondary benefit: they learn what the system does well and badly, which makes them better users of it.

Outsourcing grading to people without the domain context produces grades that measure fluency, which is the quality least in need of measurement.

How FISTA Solutions helps

FISTA Solutions builds evaluation sets incrementally from real failures rather than comprehensively up front, automates checkable criteria and reserves expert grading for judgement, validates model judges per task before relying on them, and budgets set refresh as recurring work, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.

To measure AI quality without evaluation becoming the project, message FISTA on WhatsApp, or read what is continuous evaluation.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01Why does human judgement dominate?

Because someone has to decide what good looks like and whether each output meets it. Running a model against a set is cheap; building the set, defining the criteria, and grading outputs against them is expert time and it is the bulk of the cost.

02What makes set construction expensive?

Selecting representative cases, determining the correct answer for each, and covering the failure modes including the ones where the system should abstain. That last category is the most valuable and the most often omitted because it requires deliberate thought.

03When do model judges reduce cost safely?

On criteria with clear answers — is a citation present, does the output match a schema, is a required element included. On subjective judgements they carry biases that must be measured against human grades before the judge can be trusted.

04Why do evaluation sets decay?

Because traffic changes. New products, new user populations, and new question types all appear, and a set assembled months ago describes historical usage. Refreshing it from production traffic is recurring work with a recurring cost.

05Why is no evaluation the most expensive option?

Because it means shipping changes blind. Regressions reach users, are discovered by complaint, and are diagnosed without a baseline to compare against. That cost is diffuse, which is why it is consistently under-weighted.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project