FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Trends ¡ 5 minute read

Why Data Quality Decides AI Outcomes More Than Model Choice

AI systems inherit the quality of the data they read, and most organisations have worse data than they believe. Duplicates, contradictions, stale records, and rules that exist only in people's heads cause more production failures than any model limitation, and no model choice compensates for them.

By FISTA Solutions¡ AI-Native Engineering Team¡
Why Data Quality Decides AI Outcomes More Than Model Choice article cover

Model choice gets the attention; data quality decides the outcome. This piece covers why, and what to do about it, drawing on FISTA Solutions' AI enablement delivery work.

What actually goes wrong?

The failure modes, in order of frequency.

ProblemHow it presents
Stale documentsConfidently outdated answers
Contradictory sourcesAnswers that vary by phrasing
DuplicatesOne version retrieved, others ignored
Undocumented exceptionsRules never applied
Missing metadataCannot filter by date or scope
Inconsistent structurePoor chunking, lost context

Why is staleness the worst offender?

Because nothing marks it.

A document that was accurate two years ago and is wrong now looks identical to a current one. Retrieval finds it, the model uses it, and the answer is delivered with the same confidence as a correct one.

The fix is metadata and process: dates on everything, ownership for every document, review cycles, and retrieval that can prefer or filter by recency. Most knowledge bases have none of this. See AI knowledge base quality checklist.

Why do contradictions matter more than gaps?

Because a gap is recoverable and a contradiction is not visible.

When material is missing, a well-built system says it does not know, and a person handles it. That is a manageable outcome.

When two sources disagree, retrieval returns whichever ranked higher and the model answers from it. The user sees one confident answer and no indication that the organisation holds two positions. The same question, phrased differently, may return the other.

What does duplication do?

It splits relevance and hides the authoritative version.

Five near-identical copies of a policy, with one updated, means retrieval may return any of them. It also dilutes ranking, because the signal is spread across duplicates rather than concentrated.

Deduplication before indexing, with a clear authoritative source per topic, is one of the highest-return data tasks and one of the least glamorous.

Why do undocumented rules break systems?

Because a system can only apply what it can read.

Experienced people carry exceptions — this customer is handled differently, this product has a special rule, this situation always escalates. None of it is written down, and all of it is essential to correct handling.

Extracting that knowledge is a substantial piece of work involving the people who hold it. It is also where a lot of the value is, because writing those rules down improves the human process too. See forward deployed engineer knowledge transfer.

How should quality be measured?

At the source, not only at the output.

Measure document age distribution, duplicate rate, coverage of the topics users actually ask about, and how often retrieval returns contradictory material. Those metrics predict answer quality better than answer quality predicts itself.

A quality dashboard for the corpus is a small piece of engineering with disproportionate value, because it turns a vague sense that answers are getting worse into a specific finding. See how to improve RAG accuracy.

Who owns this?

The business function that owns the content, with engineering providing the measurement.

Engineering cannot decide which of two contradictory policies is correct, and should not. What it can do is surface the contradiction, the staleness, and the gaps, so the owner can resolve them.

Without a named content owner, quality decays continuously and the AI system degrades with it. That ownership is a prerequisite, not a follow-up task.

What is the counter-argument?

The counter is that models are increasingly able to reason about conflicting sources and flag uncertainty, which is true and improving. It reduces the severity of the problem without removing it: a model can say two sources disagree, but it cannot know which is correct, and the organisation still has to decide.

What does this change for engineering teams?

It means data pipeline work — ingestion, deduplication, metadata extraction, refresh — is the majority of many AI projects, and should be scoped as such.

It also means building the corpus quality measurement early, because it is what tells you whether the rest of the system can possibly work.

What does this change for buyers?

It means asking vendors what happens when your data is contradictory or stale, and whether their system surfaces that or hides it.

And budgeting for your own data remediation, which no vendor can do for you because it requires decisions only you can make.

What should leaders do about it now?

Assess your corpus before funding the AI system. Age distribution, duplication, and coverage of common questions are quick to measure and will tell you whether the project is viable.

Then assign content ownership. Without it, quality decays and the system decays with it.

Does this change with agents?

It becomes more consequential, because an agent acts on what it reads rather than presenting it for review.

A wrong answer from stale data is an inconvenience. An action taken on stale data changes a record or contacts a customer incorrectly, and the review step that would have caught it is gone. See the shift from chatbots to agents.

How will you know if this is happening?

Watch for answers that vary with phrasing, for users correcting the system on facts, and for complaints that it was right last month. All three indicate corpus problems rather than model problems.

How FISTA Solutions reads this

FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: corpus quality measured as a first-class metric — age, duplication, coverage, contradiction — with named content owners responsible for resolving what it surfaces, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.

To discuss what this means for your roadmap, message FISTA on WhatsApp, or read how to improve RAG accuracy.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01Why does data quality dominate?

Because a model answers from what it is given. If the source is wrong, outdated, or contradictory, the answer will be too — fluently and confidently, which is worse than an obvious error.

02What is the most common problem?

Staleness. Documents that were correct when written and are now wrong, with nothing indicating which is which. The system retrieves them faithfully and produces an answer that used to be right.

03Why are contradictions worse than gaps?

Because a gap produces an admission of not knowing, which is recoverable. A contradiction produces a confident answer drawn from whichever source ranked higher, with no signal that another source disagreed.

04What about undocumented rules?

They cannot be retrieved, so the system cannot apply them. Every exception that lives in an experienced person's head is a case the system will get wrong until it is written down.

05How much of a project is data work?

Frequently more than half. Teams that budget for model integration and not for data remediation find the project stalls at the point where quality problems become visible.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project