Checklist ┬╖ 5 minute read
RAG Quality Checklist: Auditing a Retrieval System
Most retrieval problems present as model problems. Audit the corpus first, then chunking, then retrieval measured on its own, then context assembly, then grounding. Working in that order finds the fault faster than adjusting prompts, which is where teams usually start and rarely finish.
Most retrieval problems present as model problems, which is why teams spend weeks on prompts and find nothing. This checklist audits in the order that finds faults fastest, drawn from FISTA Solutions' AI agents production work.
What should be audited, in what order?
Six layers, checked from the bottom up.
| Layer | What failure there looks like |
|---|---|
| Corpus health | Confident answers from stale sources |
| Chunking | Right document, wrong fragment |
| Retrieval | Answer exists but was not found |
| Context assembly | Retrieved but not used |
| Grounding | Answers that cannot be checked |
| Freshness | Correct last quarter, wrong now |
Corpus health
The ceiling on answer quality is set here. See why data quality decides AI outcomes.
- Document count and age distribution measured
- Duplicates identified and an authoritative version designated
- Known contradictions listed and assigned to an owner
- Coverage checked against the questions users actually ask
- Every document has a date and an owner
- Documents with no owner are removed or assigned
- Access permissions on source documents are reflected in retrieval
Chunking and indexing
Chunking decides what can be retrieved at all. A concept split across a boundary is effectively absent.
- Chunk size chosen from content structure, not a default
- Overlap sufficient that concepts are not split
- Document structure preserved: headings, sections, tables
- Tables and lists handled deliberately rather than flattened
- Metadata attached to each chunk: source, date, section, permissions
- Chunks inspected manually for a sample of documents
- Re-chunking is possible without rebuilding everything else
Retrieval measurement
This is the measurement most systems lack and the one that attributes failures correctly.
- Each evaluation case records whether the needed passage was retrieved
- Retrieval hit rate reported separately from answer accuracy
- Results reviewed for queries where nothing relevant was returned
- Hybrid retrieval considered where exact terms matter
- Reranking evaluated rather than assumed beneficial
- The number of passages retrieved has been tuned with measurement
- Retrieval latency measured at the tail, not the average
Context assembly
Retrieved is not the same as used. See why context beats prompting.
- Passages ordered by relevance, most important first
- Source boundaries marked clearly in the assembled context
- Each passage labelled with its source and date
- Total context length bounded and monitored
- Irrelevant passages excluded rather than included defensively
- Instructions kept separate from retrieved content
- Retrieved content treated as data, never as instruction
Grounding and citation
Users who cannot verify an answer cannot catch an error, which makes citation a control rather than a courtesy.
- Every factual claim traceable to a retrieved passage
- Citations link to a specific passage, not just a document
- The system states when it cannot answer from the corpus
- Answers not supported by retrieved content are detected
- Contradictions between sources are surfaced rather than resolved silently
- Citation accuracy is itself evaluated
- Users have a way to report an incorrect citation
Freshness and maintenance
A corpus without a refresh pipeline decays continuously and silently.
- A refresh pipeline exists and runs on a schedule
- Deleted source documents are removed from the index
- Index lag behind the source is monitored and alerted
- Document age is available to retrieval for filtering or weighting
- Content owners have a review cadence
- Corpus quality metrics are reported alongside answer quality
- Re-embedding after a model change is planned and tested
What are the most common failures?
Tuning prompts against retrieval failures. Default chunk sizes. No retrieval measurement. Citations to document names rather than passages. And an index loaded once with no refresh.
Who should own this?
Retrieval engineering owns the pipeline and the measurement; the business function owning the content owns corpus quality. Splitting it any other way produces an engineering team unable to resolve contradictions it has no authority over.
How often should it run?
Full audit quarterly, corpus health metrics continuously, and retrieval measurement on every evaluation run. Re-audit after any change to chunking, embeddings, or the retrieval model.
What evidence should it produce?
Corpus health metrics with dates, retrieval hit rate over time, a sample of manually inspected chunks, and citation accuracy results. Together these show whether the system's foundation is sound.
What if everything checks out and quality is still poor?
Then the problem genuinely is generation, and the work moves to prompt structure, output validation, or model choice. That is a legitimate place to arrive, and arriving there with retrieval measured is very different from assuming it.
In practice most audits find corpus or retrieval issues well before this point. See how to improve RAG accuracy.
What should you do first?
Add retrieval hit rate to your evaluation suite. One measurement, and it tells you which half of the system to work on.
How FISTA Solutions helps
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: retrieval hit rate measured separately from answer quality, and corpus health tracked as a first-class metric with named content owners, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To adapt this checklist to your environment, message FISTA on WhatsApp, or read how to improve RAG accuracy.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Where should an audit start?
The corpus. Age distribution, duplication, contradictions, and coverage of the questions users actually ask determine the ceiling on answer quality, and no retrieval tuning raises it.
02How do you measure retrieval alone?
For each evaluation case, record whether the passage containing the answer was in the retrieved set. That single measurement splits failures into retrieval problems and generation problems.
03What goes wrong with chunking?
Chunks that split a concept across boundaries, or that are so large that relevance is diluted. Both make the right passage unretrievable regardless of how good the search is.
04Why is freshness so often missed?
Because the initial corpus load works and nobody builds the refresh. Documents change, the index does not, and answers slowly become wrong with no error anywhere.
05What makes grounding real?
Citations that point to a specific passage a user can open and check. A reference to a document name is not verifiable and does not prevent the failure it is meant to catch.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.