FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Checklist · 4 minute read

AI Knowledge Base Quality Checklist: Auditing a Corpus

The corpus sets the ceiling on answer quality, and most corpora are worse than their owners believe. Audit coverage against real questions, age distribution, contradictions, duplication, and ownership. Every gap found here is a wrong answer the model will deliver confidently.

By FISTA Solutions· AI-Native Engineering Team·
AI Knowledge Base Quality Checklist: Auditing a Corpus article cover

The corpus sets the ceiling on answer quality, and it is where audits should start rather than end. This checklist covers it, drawn from FISTA Solutions' AI enablement delivery work.

What determines corpus quality?

Six properties, all measurable.

PropertyHow to measure it
CoverageReal questions versus content
FreshnessAge distribution and refresh rate
ConsistencyContradiction count
UniquenessDuplicate rate
OwnershipDocuments with a named owner
StructureChunk quality on inspection

Coverage

Measured against what people ask, not against a topic list.

  • Questions from production logs collected and categorised
  • Questions the corpus cannot answer identified
  • Topics with thin coverage listed
  • Gaps prioritised by question frequency
  • Content commissioned for the top gaps with owners assigned
  • Coverage re-measured after content is added
  • Questions that should be out of scope identified and handled

Freshness

Continuous decay, so measure the distribution rather than the newest item.

  • Age distribution measured across the corpus
  • Documents past a defined age flagged for review
  • Review cadence agreed per content type
  • Dates visible to retrieval for filtering or weighting
  • Refresh pipeline confirmed running and monitored
  • Deleted source documents removed from the index
  • Superseded documents archived rather than left alongside current ones

Contradictions

The failure that produces confident wrong answers.

  • Known contradictory topics identified
  • Authoritative source designated per topic
  • Conflicting documents resolved or clearly superseded
  • A process exists for reporting a newly found contradiction
  • Retrieval surfaces disagreement rather than silently picking one
  • Policy and procedure documents checked against each other
  • Regional or departmental variants labelled rather than merged

Duplication

Splits relevance and hides the current version.

  • Duplicate and near-duplicate rate measured
  • Authoritative copy designated where duplicates exist
  • Copies removed or excluded from the index
  • Deduplication applied before indexing
  • Boilerplate that appears in every document handled
  • Versions of the same document consolidated
  • Ongoing deduplication built into the ingestion pipeline

Ownership

The mechanism by which decay is corrected.

  • Every document has a named owner
  • Owners are current employees in relevant roles
  • Owners know they own it and what that entails
  • Review cadence agreed with each owner
  • Documents with no owner resolved or removed
  • Ownership survives role changes through a defined process
  • Owner engagement tracked, not assumed

Structure

Determines whether good content can be retrieved. See RAG quality checklist.

  • Documents have headings and clear sections
  • Metadata present: date, owner, scope, permissions
  • Tables and lists preserved rather than flattened
  • Chunking inspected manually on a sample
  • Concepts not split across chunk boundaries
  • Very long and very short documents handled appropriately
  • File formats that convert badly identified and remediated

What are the most common failures?

Measuring coverage against a topic list. Checking freshness by looking at recent additions. Leaving contradictions for the model to resolve. Indexing duplicates. And documents with no owner, which is where decay concentrates.

Who should own this?

The business function that owns the content owns its quality. Engineering supplies measurement and tooling. Corpus quality owned by engineering produces metrics nobody acts on.

How often should it run?

Full audit quarterly with continuous metrics on age and coverage. Re-audit after any bulk content addition, which is when duplication and contradictions usually enter.

What evidence should it produce?

Corpus metrics over time, the gap list with owners, contradiction resolutions, and the ownership coverage percentage. Those show whether quality is being maintained or merely observed.

What if the corpus is too large to audit fully?

Sample, and prioritise by usage. The documents retrieved most often matter most, and auditing the top few hundred covers the majority of answers.

Retrieval logs tell you which documents are actually used. Content nobody retrieves can be audited later or removed, and removal is frequently the right answer. See why data quality decides AI outcomes.

What should you do first?

Pull the hundred most-retrieved documents and check their dates. The age distribution of what is actually used tells you more than the corpus average.

How FISTA Solutions helps

FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: corpus quality measured against the questions users actually ask, with named content owners and freshness tracked on the documents retrieval actually uses, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.

To adapt this checklist to your environment, message FISTA on WhatsApp, or read RAG quality checklist.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01How do you measure coverage?

Against the questions users actually ask, taken from logs, rather than against a topic list. The gap between what people ask and what the corpus covers is the real coverage problem.

02What does stale content do?

Produces answers that were correct once. The system retrieves faithfully and the user receives outdated information with full confidence and no warning.

03Why are contradictions worse than gaps?

Because a gap produces an admission of not knowing. A contradiction produces a confident answer from whichever source ranked higher, with no indication another source disagreed.

04What does ownership change?

Whether decay gets corrected. An owned document is reviewed and updated; an unowned one sits until someone notices it is wrong, usually via a customer.

05How does structure affect retrieval?

Documents without headings, dates, or clear sections chunk poorly, which makes the right passage hard to retrieve even when the content is correct.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project