FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Playbook ¡ 6 minute read

How to Overhaul a Knowledge Base So AI Can Use It

A knowledge base becomes usable for AI retrieval when every document has an owner and a review date, contradictions are resolved rather than left to coexist, structure supports clean chunking, and stale content is excluded from retrieval rather than silently served alongside what is current.

By FISTA Solutions¡ AI-Native Engineering Team¡
How to Overhaul a Knowledge Base So AI Can Use It article cover

Retrieval systems inherit every flaw in the knowledge base behind them. Contradictions, stale policies, and unowned documents all become confidently wrong answers. This playbook covers fixing that, drawing on FISTA Solutions' AI enablement work.

When is this worth doing?

When a retrieval system is planned or already performing badly, and the cause is content rather than technique.

The diagnostic is simple: read the documents retrieval surfaced for a failed question. If the right answer was not in the corpus, or was contradicted elsewhere in it, the problem is content and no amount of embedding or reranking work will fix it.

What does the sequence look like?

StepPurpose
1. Map and scopeWhich content will answer which questions
2. Assign ownersEvery document, a named person
3. Resolve contradictionsDecide which version is authoritative
4. Restructure for chunkingHeadings and self-contained sections
5. Exclude the staleReview dates enforced by the system
6. Measure answer qualityAgainst real questions, not corpus size

Step 1 — Map the content against real questions

Start from the questions people actually ask rather than from the document library. Collect real questions from support tickets, internal channels, or a sample of searches, and check which documents would answer them.

That mapping produces three lists: content that answers questions well, content that answers nothing anyone asks, and questions nothing answers. All three are useful, and the third is usually the most valuable.

Scope the overhaul to the content that matters. A knowledge base has a long tail nobody reads, and cleaning it is effort with no return.

Step 2 — Assign an owner to every document

Every document that will inform answers needs a named person responsible for whether it is correct.

This is the step organisations resist and the one that makes everything else sustainable. Unowned documents are how corpora decay: nobody updates them, nobody retires them, and nobody notices when they contradict something newer.

Where no owner can be found, that is itself the answer — content nobody will own should not be informing answers, and removing it is usually uncontroversial once framed that way.

Step 3 — Resolve contradictions

Find documents that say different things about the same subject and decide which is authoritative.

This is a business decision, not a technical one, and it requires the owner. A retrieval system faced with two contradictory sources will surface one of them, and which one depends on phrasing rather than on correctness.

Contradictions are more common than teams expect, particularly around policies that changed. The old version rarely gets removed; it just stops being the one people refer to, which works for humans who know the history and fails completely for a system that does not.

Step 4 — Restructure for chunking

Documents get split before they are indexed, and how they are written determines whether those chunks make sense alone.

Clear headings, self-contained sections, and explicit subjects help. A section beginning 'This applies in the following cases' produces a chunk that has lost what 'this' refers to, and the system will answer from it anyway.

The practical fix is modest editing: making headings descriptive, repeating the subject at the start of sections, and breaking long undifferentiated text into labelled parts. See what is chunk overlap.

Step 5 — Exclude stale content automatically

Give every document a review date and exclude anything past it from retrieval until an owner confirms it.

That automatic exclusion is what creates the pressure to maintain. A policy asking owners to review annually produces nothing; a system that stops serving their content until they do produces reviews.

Expect resistance initially and a rapid improvement in currency. The first cycle removes a surprising proportion of the corpus, and most of what disappears turns out not to be missed.

Step 6 — Measure answer quality against real questions

Build an evaluation set from the real questions gathered in step one, with correct answers agreed by the owners, and measure the system against it.

Corpus size, document count, and index freshness are activity metrics. Whether a real question gets a correct, current, well-sourced answer is the outcome.

Run it regularly. Knowledge bases decay continuously, and a measurement taken once tells you about the day it ran. See what is continuous evaluation.

What about content in other systems?

Most organisations have knowledge scattered across a wiki, a ticketing system, shared drives, and people's inboxes.

Deciding what is in scope matters more than reaching everything. A retrieval system over a curated, owned, current subset outperforms one over everything, because the everything version surfaces drafts, superseded versions, and personal notes with equal confidence.

Start narrow and expand once the controls are working.

What about access control?

Retrieval must respect it. A system that indexes everything and answers from whatever is relevant will surface content the asker should not see, and that failure is both a security incident and a trust event.

Enforce permissions at retrieval time against the asking user rather than filtering afterwards, and test it deliberately with users at different permission levels. See what is metadata filtering in rag.

Who needs to be involved?

A programme owner, the document owners from each function, and an engineer who can measure retrieval quality.

The document owners are the constraint. Without their participation, the overhaul produces a cleaner snapshot that decays at the same rate as before.

How long does it take?

Four to eight weeks for a scoped overhaul of the content that matters, then continuous. The review-date mechanism becomes business as usual rather than a project.

What are the common failure modes?

Cleaning the whole corpus rather than what answers real questions. Leaving contradictions in place. No owners. Ignoring chunking. Serving stale content and hoping ranking handles it. And measuring corpus size.

How do you know it worked?

Real questions answered correctly and currently, contradictions resolved, every retrievable document owned and within its review date, and quality holding steady over months rather than decaying.

What does it cost?

Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.

Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.

What should you do first?

Take twenty real questions people asked last week and check what your current corpus would answer. The failures tell you exactly what to fix.

How FISTA Solutions helps

FISTA Solutions runs this work alongside client teams rather than around them: content scoped to the questions people actually ask, review dates enforced by the system so currency does not depend on goodwill, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.

To run this with support, message FISTA on WhatsApp, or read how to improve RAG answer quality.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01Why does knowledge base quality matter so much?

Because a retrieval system inherits everything wrong with it. Contradictions produce inconsistent answers, stale documents produce confidently wrong ones, and unowned content means nobody notices either until a user does.

02What is the most damaging defect?

Contradiction. Two documents stating different things about the same policy produce answers that vary by which one retrieval happened to surface, which destroys trust faster than an obvious failure would.

03How does structure affect retrieval?

Through chunking. Documents with clear headings and self-contained sections chunk cleanly; wall-of-text documents produce chunks that lose their context, and the system then answers from fragments that no longer mean what they did.

04Should stale content be deleted or excluded?

Excluded from retrieval, at minimum, with deletion decided separately. A document past its review date should not be informing answers, even if it needs retaining for other reasons.

05How do you keep it from decaying again?

Owners and review dates enforced by the system rather than the policy. Content past its date gets flagged and excluded automatically, which creates the pressure that keeps owners engaged.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project