FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Whitepaper · 8 minute read

RAG Evaluation Methodology: An Engineering Whitepaper

RAG evaluation must measure retrieval and generation separately, because most quality failures are retrieval failures that generation metrics cannot diagnose. The method is a reference set of real questions with known required facts, retrieval metrics against those facts, groundedness and completeness scoring on answers, and continuous running rather than release-time testing.

By FISTA Solutions· AI-Native Engineering Team·
RAG Evaluation Methodology: An Engineering Whitepaper article cover

Retrieval-augmented generation is the dominant enterprise AI architecture and the most poorly evaluated. The typical practice is to ask a handful of questions during development, read the answers, declare the system good, and ship it, after which quality degrades as content changes and nobody notices until users stop trusting it. The reason is not laziness but a missing method: teams know how to test deterministic software and have no framework for a system whose output is probabilistic and whose inputs change continuously. This whitepaper provides that framework. It draws on FISTA Solutions' AI enablement practice and complements the AI evaluation and testing whitepaper and the enterprise RAG reference architecture whitepaper.

Why separate retrieval from generation?

Because they are different systems with different failure modes and unrelated fixes, and an end-to-end score conflates them.

Consider an answer that is wrong. If the passage containing the correct information was never retrieved, the generation model did the best it could with what it had, and every hour spent on prompt engineering is wasted. If the passage was retrieved and the model still answered wrongly, retrieval tuning will not help.

Separating the measurement localises every failure immediately:

RetrievalGenerationDiagnosis
Facts presentAnswer correctWorking
Facts absentAnswer wrongRetrieval, chunking, or corpus gap
Facts absentAnswer confidently wrongAlso a grounding and abstention failure
Facts presentAnswer wrongPrompt, model, or context assembly
Facts presentAnswer incompleteContext ordering or output constraints
Facts partially presentAnswer plausible but partialRetrieval depth or completeness

Teams that adopt this split typically discover that the majority of their failures are in the first column, which redirects effort from model selection to content and retrieval work.

How is a reference set built?

From real questions, which is the point most teams get wrong by inventing questions that resemble what they imagine users ask.

Sources: support tickets and their resolutions, internal search queries including those returning nothing, recorded user sessions, the questions the business team says are most common, and the failures the system has already produced.

For each case, record the question as actually phrased; the correct answer; the specific documents and passages containing the information; and any acceptable variations. Include three categories teams routinely omit: questions whose correct answer is that the corpus does not contain the information, which tests abstention; questions with multiple valid answers depending on context, which tests clarification; and questions that are ambiguous, which tests whether the system asks rather than guesses.

Size matters less than coverage. A few hundred well-chosen cases spanning the question types, document types, and edge cases beats thousands of similar easy ones. See what is a golden dataset.

What retrieval metrics matter?

Not the information retrieval textbook set, which assumes a different problem. The practical metrics are:

Fact recall at k: the proportion of reference cases where the required passages appear in the top k retrieved results. This is the primary metric, because a fact not retrieved cannot be used.

Precision within the returned set: how much of what was retrieved is relevant, which matters because irrelevant context degrades generation and costs tokens.

Rank of the required passage: whether the correct passage arrived first or eighth, since position affects how the model weights it.

Coverage: the proportion of questions for which the corpus contains an answer at all, which distinguishes a retrieval problem from a content gap and points at different remediation.

Segment all of these by question type and document type, because aggregate retrieval quality hides the document class that parses badly.

How is generation scored?

Against three properties, measured separately.

Groundedness. Every claim in the answer is supported by the retrieved context. Scored by decomposing the answer into discrete claims and checking each against the cited passages. This is the property that matters most in enterprise settings, because an ungrounded claim is a fabrication delivered with the organisation's authority.

Correctness. The answer matches the reference answer in substance. Distinct from groundedness: an answer can be grounded in a retrieved passage that was itself the wrong passage.

Completeness. The answer contains what the question required, not merely something true. This is the most commonly missed failure, because partial answers read as satisfactory and only fail when the user acts on them.

Add abstention correctness: did the system decline when it should have, and not decline when it should not have. Systems tuned only for correctness learn to answer everything.

How is scoring done at scale?

With a model as judge, calibrated against human scoring. The practical method scores a sample of cases by human experts first, runs the model judge over the same sample, measures agreement, and refines the rubric until agreement is acceptable for the stakes involved. Only then is the judge used at scale, with a continuing human-scored sample to detect judge drift.

Rubrics must be specific. A judge asked whether an answer is good produces noise; a judge asked whether each claim is supported by the cited passage, answering per claim, produces usable signal. See what is llm as a judge.

What should be monitored in production?

Leading indicators, because quality metrics lag. The most useful are:

Retrieval score distributions. When the top result's similarity score distribution shifts downward, retrieval has degraded, usually because content changed or the corpus grew. This moves before answer quality visibly falls.

No-result rate. Queries returning nothing above threshold indicate coverage gaps, and the questions themselves are the content roadmap.

Citation rate and distribution. If a small number of documents account for most citations, the corpus may be narrower than assumed; if citations spread thinly, retrieval may be unfocused.

User signals. Corrections, rephrasings, escalations, and abandonment, each of which indicates failure without anyone filing a report.

Sampled live scoring. A small proportion of real traffic scored against the same rubric, which is the only direct measure of production quality.

When does evaluation run?

On every change to anything that affects behaviour: corpus content, chunking strategy, embedding model, retrieval configuration and thresholds, reranking, context assembly, prompts, and generation model. Plus a scheduled run, because the corpus changes continuously through edits nobody flags as a system change.

The gate should be quality, not absence of errors. A change that improves one question type and degrades another is a trade-off requiring a decision, which is visible only when evaluation is segmented.

How are content gaps distinguished from system faults?

By the coverage metric. If the required information is not in the corpus, no amount of retrieval tuning helps, and the remediation is content work with an owner rather than engineering.

This distinction matters organisationally. Engineering teams asked to fix quality problems caused by missing or outdated content will spend months on retrieval configuration without improvement. Reporting coverage separately makes the content owner's accountability visible. See the enterprise knowledge management whitepaper.

What does a mature evaluation setup look like?

A versioned reference set with an owner, growing from production failures and shrinking as cases become obsolete. Retrieval and generation scored separately and segmented by question and document type. A calibrated judge with ongoing human validation. Evaluation wired into the deployment pipeline as a gate. Production monitoring on leading indicators with alerting. And a regular review where retrieval, generation, and coverage trends are examined together with the content owners present.

What goes wrong?

Reference sets of invented questions. End-to-end scores only. Groundedness unmeasured, so fabrication goes undetected. Completeness ignored. No abstention cases, so the system learns to always answer. Judges uncalibrated. Evaluation at release only. Retrieval score distributions unmonitored. And content gaps attributed to the engineering team, which guarantees they are never fixed.

How does this scale across multiple RAG systems?

Through shared infrastructure and per-system reference sets. The harness, judge, rubrics, and monitoring are common; the reference sets and thresholds are specific. Organisations running several retrieval systems without this end up with inconsistent quality claims and no way to compare, which makes portfolio decisions arbitrary.

What does a worked example look like?

Take a support assistant answering product questions. The reference set holds four hundred real questions from tickets, each with the correct answer and the specific documentation sections containing it, including forty questions the corpus cannot answer and thirty that are ambiguous.

Retrieval evaluation reports fact recall at five of, say, eighty-two percent overall, but segmented it shows ninety-one percent for questions answered by the product manual and sixty-three percent for those answered by release notes. That single segmentation points directly at the fix: release notes are chunked badly because their structure differs from the manual.

Generation evaluation on the cases where facts were retrieved reports groundedness of ninety-six percent, correctness of ninety-four percent, and completeness of eighty-one percent. Completeness is the weak point, and inspection shows the model answers the first part of multi-part questions and stops, which is a prompt and output-constraint issue rather than a retrieval one.

Abstention correctness on the forty unanswerable questions is sixty percent, meaning the system fabricates an answer two times in five when it should decline. That is the most serious finding, and it would be invisible in any evaluation that only asked answerable questions.

Four measurements, three distinct fixes, none of which is changing the model. That is what the methodology buys.

How FISTA Solutions delivers this

FISTA Solutions builds RAG systems with separated retrieval and generation evaluation, reference sets drawn from real questions, calibrated groundedness scoring, production monitoring on leading indicators, and evaluation gates in the deployment pipeline, so quality is measured rather than assumed, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.

To find out whether your retrieval system actually works, message FISTA on WhatsApp, or read the AI evaluation and testing whitepaper.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01Why separate retrieval and generation evaluation?

Because they fail differently and the fixes are unrelated. If the needed facts were not retrieved, no prompt or model change helps; if they were retrieved and the answer is still wrong, no retrieval tuning helps. End-to-end scores alone tell you something is wrong without telling you what.

02How do you build a RAG reference set?

From real questions, drawn from support tickets, search logs, and user sessions, with the correct answer and the specific documents or passages that contain it recorded for each. Include the questions the system currently fails, and the ones whose correct answer is that the corpus does not contain it.

03What is groundedness and how is it scored?

Whether every claim in an answer is supported by the retrieved context. It is scored by decomposing the answer into claims and checking each against the cited passages, using a model as judge calibrated against human scoring on a sample, with disagreements reviewed.

04What should be monitored in production?

Retrieval score distributions, which drift before answers degrade; the rate of queries returning nothing above threshold; citation rates; user corrections and escalations; and periodic sampled scoring of live traffic against the same rubric used in offline evaluation.

05How often should RAG evaluation run?

On every change to the corpus, chunking strategy, embedding model, retrieval configuration, reranking, prompts, or generation model, plus a scheduled run to catch drift from content edits nobody flagged as a system change. Release-time-only evaluation misses most real regressions, because the inputs change continuously while the code does not.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project