Playbook · 6 minute read
How to Improve RAG Answer Quality Systematically
Improving retrieval-augmented answers starts with diagnosing which layer failed: whether the right content was in the corpus, whether retrieval found it, and whether the model used it correctly. Most quality problems are retrieval problems, and prompt work cannot fix them.
Most retrieval-augmented quality problems are retrieval problems that look like generation problems, which is why prompt tuning rarely fixes them. This playbook covers diagnosing the layer and fixing the right one, drawing on FISTA Solutions' AI enablement work.
When is this worth doing?
When a retrieval system gives wrong, incomplete, or inconsistent answers, and before anyone starts adjusting prompts in response to individual complaints.
The symptom that indicates this playbook rather than another is inconsistency: the same question answered differently on different days, or a correct answer for one phrasing and a wrong one for another.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Assemble failed questions | Real ones, with expected answers |
| 2. Diagnose the layer | Corpus, retrieval, or generation |
| 3. Fix the corpus | Missing, stale, contradictory content |
| 4. Fix retrieval | Chunking, embedding, filtering, reranking |
| 5. Fix generation | Prompting, citations, abstention |
| 6. Measure per layer | So the next regression is attributable |
Step 1 — Assemble real failed questions
Collect twenty to fifty questions the system answered badly, with what the correct answer should have been, agreed by someone who knows.
This set is the instrument for everything that follows. Without it, improvement is a sequence of changes with no way to tell whether any of them helped.
Draw from real usage rather than inventing questions. Real users ask things in ways nobody anticipates, and those phrasings are where retrieval breaks.
Step 2 — Diagnose the layer for each failure
For each failed question, check three things in order. Is the correct answer present in the corpus at all. Did retrieval return the passage containing it. Did the model use that passage correctly.
That sequence attributes the failure to a layer, and the layers have entirely different fixes. A missing document is a content problem; a document present but not retrieved is a retrieval problem; a retrieved document ignored is a generation problem.
Expect the distribution to be heavily weighted towards the first two. Teams that skip this step and adjust prompts are fixing the layer least likely to be at fault.
Step 3 — Fix the corpus first
Content that is missing, stale, or contradictory cannot be fixed downstream.
Missing content needs writing or ingesting. Stale content needs excluding — a superseded policy retrieved confidently is worse than no answer. Contradictions need resolving by an owner who can say which version is authoritative.
This is frequently the largest single improvement and the least technical. See how to overhaul a knowledge base for ai.
Step 4 — Fix retrieval systematically
Work through the retrieval layer in order of expected return: chunking, then filtering, then reranking, then embedding choice.
Chunking first, because it has the largest effect and the least attention. Chunks that split a concept, or that lose the context of their document, produce passages the model cannot use even when they are returned. Overlap and heading-aware splitting both help.
Metadata filtering removes candidates that should never have been considered — wrong product, wrong region, wrong date. Reranking a wider candidate set with a stronger model improves what reaches generation, usually noticeably. See what is sparse vs dense retrieval.
Step 5 — Then fix generation
Once the right passages are arriving, work on how they are used: requiring citations, instructing the model to answer only from the provided context, and enabling explicit abstention when the context does not contain the answer.
Abstention is the change that most improves trust. A system that says it does not know is more useful than one that produces a plausible answer from insufficient context, because the second is indistinguishable from a correct answer to the reader. See what is abstention in ai.
Citations serve double duty: they let readers verify and they let you diagnose future failures precisely.
Step 6 — Measure per layer
Track retrieval quality — whether the correct passage was in the returned set — separately from answer quality.
End-to-end measurement tells you the system got worse without telling you where. Per-layer measurement makes the next regression attributable in minutes rather than days.
Run both on every corpus change as well as every code change. A retrieval system whose knowledge base changed has changed, and evaluation triggered only by deployments will miss it.
What about query rewriting?
It helps where users ask questions in terms the corpus does not use, which is common in technical domains with internal vocabulary.
Rewriting the query before retrieval — expanding acronyms, adding synonyms, or generating several variants — improves recall meaningfully. It also adds latency and a failure mode of its own, so measure the gain rather than adopting it because it sounds sophisticated. See what is query expansion.
When is the answer not RAG at all?
When the question requires computation, aggregation, or current data from a system of record rather than retrieval from documents.
Asking a retrieval system how many open cases a customer has is asking the wrong architecture. That is a query against a database, and dressing it as retrieval produces answers assembled from whatever documents mentioned similar things.
Route those questions to tools rather than to retrieval, and be explicit about which questions the system does not handle.
Who needs to be involved?
Someone who can judge correct answers in the domain, an engineer who can change retrieval, and an owner for the corpus.
The domain judge is essential. Retrieval improvements evaluated by engineers who cannot tell a correct answer from a plausible one optimise for plausibility.
How long does it take?
Two to four weeks for a systematic pass, with the corpus work usually being the longest and the retrieval changes the quickest to test.
What are the common failure modes?
Adjusting prompts without diagnosing. Ignoring corpus problems. Changing several retrieval parameters at once. No abstention path. And measuring only end to end.
How do you know it worked?
Failure rate down on the assembled question set, retrieval recall measurably improved, answers carrying citations that check out, and the system abstaining rather than inventing when context is thin.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
Take ten failed questions and check, for each, whether the answer was in the corpus at all. That five-minute exercise usually reallocates the whole improvement effort.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: failures diagnosed by layer before anything is changed, retrieval and answer quality measured separately so regressions stay attributable, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read what is retrieval augmentation.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01How do you diagnose a RAG failure?
Take a failed question and check three things in order: was the answer in the corpus, did retrieval return it, and did the model use it correctly. Each has a different fix and they are frequently confused.
02Why are most failures retrieval failures?
Because generation is comparatively reliable when given the right context. If the correct passage was retrieved and the model still answered wrongly, that is unusual; far more often the passage was never surfaced.
03How much does chunking matter?
A great deal. Chunks that split a concept across boundaries, or that lose their context when separated from the document, produce retrieval that returns technically relevant fragments the model cannot use correctly.
04Does reranking help?
Usually, and it is one of the higher-return changes available. Retrieving a wider candidate set and reranking with a stronger model improves the passages reaching the generation step, at a modest cost.
05Why do citations matter for quality?
Because they make failures visible. An answer citing its sources lets a reader verify in seconds and lets you diagnose failures precisely, whereas an unsourced answer is either trusted or discarded with no information either way.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.