Glossary · 5 minute read
What Is an Evaluation Rubric? Grading AI Output Explained
An evaluation rubric defines what good output means as specific, independently gradable criteria, so that different graders reach the same judgement on the same output. Vague dimensions such as helpful or high quality produce scores that vary by grader and cannot support a decision.
Evaluation is only as good as the definition of good it applies, and most evaluation rubrics are written quickly with dimensions like accuracy, helpfulness, and tone that different people interpret differently. The resulting scores move for reasons nobody can trace. This explainer covers how to write criteria that produce stable measurement. It complements ai evaluation checklist and what is offline vs online evaluation, and reflects FISTA Solutions' approach in AI enablement delivery.
What makes a criterion usable?
Agreement. If two competent graders apply it to the same output and reach different verdicts, the criterion is not doing its job, and any score built on it is noise.
"Does the answer cite a source for every factual claim" is gradable â you can check. "Is the answer helpful" is not, because helpfulness depends on who asked and what they needed, neither of which the grader knows.
| Criterion style | Gradable | Note |
|---|---|---|
| Cites a source for each claim | Yes | Checkable |
| States when it cannot answer | Yes | Checkable |
| Follows the required format | Yes | Often automatable |
| Contains no unsupported claim | Yes | Requires source access |
| Is helpful | No | Depends on unknown context |
| Has appropriate tone | Weakly | Needs explicit definition |
Why prefer short scales?
Because graders cannot distinguish fine gradations reliably. The difference between a seven and an eight is grader-dependent, varies by mood and fatigue, and contributes noise that swamps the effect you are trying to measure.
Binary judgements per criterion, aggregated across several criteria, produce a more stable and more interpretable score. Knowing that a system cites sources in 94% of cases and abstains correctly in 71% tells you what to fix. Knowing it scored 7.3 does not.
How should graders be calibrated?
By grading the same sample independently and comparing. Disagreements are then discussed, and the rubric is refined until agreement is acceptable.
The important insight is that disagreement almost always indicates an ambiguous criterion rather than a poor grader. Treating it as a rubric problem rather than a people problem is what makes calibration productive, and each round improves every grade that follows.
Is model-as-judge reliable?
It depends on the criterion. For checkable properties â is a citation present, does the output match the schema, is a required element included â model judges work well and scale to volumes human grading cannot reach.
For subjective dimensions they carry known biases: preferring longer answers, preferring their own stylistic register, and being influenced by confident phrasing. Before relying on one, grade a sample with humans and measure agreement with the judge on that specific task. A judge validated on one task does not transfer to another.
How many criteria?
Three to six. Long rubrics are applied inattentively, and a criterion buried at position eleven gets the same glance as everything else regardless of how much it matters.
If more dimensions genuinely matter, split the evaluation by output type rather than extending one rubric to cover everything.
What should the rubric encode?
What actually matters for the decision the output supports. This sounds obvious and is frequently missed: rubrics get written around what is easy to grade rather than around what would make the system worth using.
A rubric for a support assistant that measures tone and format but not whether the answer resolved the question is measuring the wrong thing precisely.
How does it evolve?
Deliberately and with versioning. Changing a rubric invalidates comparison with prior results, so changes should be recorded and, where comparison matters, a sample re-graded under both versions to establish the relationship. Silent rubric drift is a common reason evaluation trends stop meaning anything.
What should you do first?
Take your current rubric and have two people grade twenty outputs independently. The disagreement rate will tell you immediately whether your criteria are doing any work, and the specific disagreements will tell you which ones to rewrite.
How does a rubric relate to the product decision?
It should be derived from it. The question a rubric exists to answer is whether this output is good enough for what the system is for, and that depends on what happens next: whether a human reviews it, whether it reaches a customer, whether an action follows automatically.
An output that is 80% right is excellent when a human reviews it and unacceptable when it triggers a payment. Rubrics written without that context tend to grade generic quality rather than fitness for the actual use, which is why two teams evaluating similar systems can reasonably arrive at very different criteria.
What about criteria that automation can check?
Move them out of human grading entirely. Format compliance, schema validity, presence of required fields, and length constraints are all checkable in code, faster and more reliably than any grader. Reserving human and model grading for judgements that genuinely need judgement keeps the expensive capacity focused where it adds something.
How FISTA Solutions helps
FISTA Solutions writes evaluation criteria that graders can apply consistently, uses short scales aggregated across criteria, calibrates graders and measures agreement, validates model judges against human grades per task, and versions rubrics so trends remain comparable, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To measure AI quality in a way that supports decisions, message FISTA on WhatsApp, or read ai evaluation checklist.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What makes a criterion good?
That two competent graders applying it to the same output reach the same verdict. Does the answer cite a source for every factual claim is gradable. Is the answer helpful is not, because helpfulness varies with who is asking and what they needed.
02Why prefer short scales?
Because graders cannot reliably distinguish a seven from an eight, and the resulting noise swamps the signal. Binary judgements per criterion, aggregated across several criteria, produce more stable and more interpretable measurement than a single broad score.
03What is grader calibration?
Having several graders score the same sample independently, comparing results, and resolving disagreements by refining the rubric. Disagreement almost always reveals an ambiguous criterion rather than an incompetent grader, and fixing it improves every future grade.
04Is model-as-judge reliable?
For criteria with clear right answers it works well and scales cheaply. For subjective dimensions it has biases â toward longer answers, toward its own style â that must be measured against human grades on the specific task before the judge is trusted.
05How many criteria should a rubric have?
Few enough to apply consistently, usually three to six. Long rubrics get applied inattentively, and a criterion buried at position eleven receives the same cursory glance as everything else regardless of how much it actually matters to the decision the output supports.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.