Playbook · 7 minute read
How to Run an AI Data Cleanup That Pays for Itself
A data cleanup pays for itself when it is scoped to a specific AI use rather than to the data estate, measured against what actually blocks that use, fixed at source where possible, and stopped when the use case works rather than when the data is perfect.
Data cleanup projects run forever when nobody defines what clean means. Scoped to a specific AI use, they have an endpoint and a measurable payoff. This playbook covers running one that finishes, drawing on FISTA Solutions' AI enablement work.
When is this worth doing?
When a specific AI use case is blocked by data problems, and somebody can say what the system needs in order to work.
It is not worth doing as preparation for unspecified future AI work. Cleaning data for a use case that does not exist yet produces effort against guesses, and the guesses are usually wrong about which fields matter.
What does the sequence look like?
| Step | Purpose |
|---|---|
| 1. Scope to the use case | What this system needs, not general quality |
| 2. Measure the actual defects | Sample real records against the requirement |
| 3. Separate fix from handle | Some defects are cheaper to tolerate |
| 4. Fix at source where possible | Otherwise the problem returns |
| 5. Backfill only what matters | History nobody reads is not worth cleaning |
| 6. Stop and document | Record what you chose not to fix |
Step 1 — Scope to the use case
Start from what the system needs rather than from the data estate. A retrieval system needs documents that are current, attributed, and deduplicated; a classification system needs consistent labels; an extraction system needs source documents it can read.
Write that requirement down as a list of specific properties. That document is the definition of clean for this project, and having it is what lets the project end.
Resist scope expansion. Every cleanup surfaces adjacent problems, and fixing them is usually reasonable and always the mechanism by which the project stops finishing.
Step 2 — Measure the actual defects
Sample real records and measure against the requirement. How many documents lack an owner, how many records have the field the system needs, how many duplicates exist, how stale the corpus is.
Sampling beats full analysis at this stage. A few hundred records give you the rates you need to plan, and the full analysis can wait until you know which defects matter.
The result is frequently surprising in both directions: a defect everyone complains about turns out to affect two per cent of records, while one nobody mentioned affects a third.
Step 3 — Separate what to fix from what to handle
For each defect, ask whether it is cheaper to fix the data or to handle it in the system.
A missing optional field can often be handled: the system treats it as unknown and behaves sensibly. A wrong value cannot, because the system will act on it confidently. That distinction drives most of the decisions.
Handling has a cost too — complexity in the system, and a permanent workaround. But it is bounded, whereas backfilling three years of records is not, and honest comparison usually favours handling more often than teams expect.
Step 4 — Fix at source where possible
Where a defect is being created continuously, fix the process creating it: validation at entry, a required field, a picklist instead of free text, or a changed workflow.
Without that, downstream cleaning becomes permanent infrastructure. A pipeline that normalises a field entered inconsistently has to run forever, and it will eventually be the only thing anyone remembers about the problem.
This step involves other teams and is the slowest part of the project. Start it early, in parallel with the rest, because it has the longest lead time and the largest payoff.
Step 5 — Backfill only what matters
Historical records need cleaning only if something will read them.
For a retrieval system, that means the documents people actually search for — usually a fraction of the corpus. For an analytics use, it means the period the analysis covers. For most uses, the deep history is not worth touching.
Decide the cutoff explicitly and record it. Undocumented partial backfills produce later confusion when someone notices the older records look different and assumes it is a bug.
Step 6 — Stop, and document what you left
When the use case performs acceptably against its evaluation set on real inputs, stop.
Then write down what you chose not to fix and why. That document prevents the next team from rediscovering the same decisions, and it prevents the same defects being reported as new findings in six months.
It also matters for governance: knowing the data's known limitations is part of understanding the system's limitations, and both belong in the system documentation. See how to document an ai system.
What about deduplication?
It matters more for retrieval than teams expect. Near-duplicate documents crowd results, making a system look worse than it is, and they multiply the cost of indexing and embedding.
Exact duplicates are easy. Near-duplicates — the same policy in three slightly different versions — require deciding which is authoritative, which is a business decision rather than a technical one and needs an owner to make it.
What about stale content?
Stale documents are the most damaging single defect in retrieval systems, because the system will cite a superseded policy with complete confidence.
The fix is ownership and dates, not cleaning. Every document in a corpus that informs answers needs an owner and a review date, and anything past its date should be flagged or excluded rather than silently served. That is an ongoing control rather than a cleanup task, and establishing it is usually the most valuable output of the project.
Who needs to be involved?
Someone who can say what the use case needs, an owner from the team that produces the data, and an engineer who can measure and transform it.
The data producer matters most. Fixes at source require their cooperation, and cleanups run without them produce downstream workarounds instead.
How long does it take?
Two to six weeks for a scoped cleanup, with source fixes continuing longer because they involve process change in other teams.
Projects running past a quarter have usually lost their scope and should be re-anchored to a specific use case.
What are the common failure modes?
Scoping to the estate rather than a use case. Measuring general quality. Cleaning downstream instead of fixing at source. Backfilling history nobody reads. And never declaring the work finished.
How do you know it worked?
The use case meeting its evaluation threshold on real inputs, defect rates at source falling rather than being corrected downstream, and a written record of what was deliberately left alone.
What does it cost?
Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.
Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.
What should you do first?
Write down the specific data properties your AI use case requires, then sample two hundred real records against that list. The result scopes the project.
How FISTA Solutions helps
FISTA Solutions runs this work alongside client teams rather than around them: cleanup scoped to what a specific system needs rather than to the estate, defects fixed at source where they are being created, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.
To run this with support, message FISTA on WhatsApp, or read AI data cleanup cost.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why do data cleanup projects never finish?
Because clean was never defined. A cleanup scoped to an estate has no endpoint, while one scoped to making a specific system work has a clear test: does the system produce acceptable results on real inputs.
02What should you measure?
The specific defects that block your use: missing fields the system needs, inconsistent codes it cannot interpret, duplicates that distort retrieval. General quality scores tell you little about whether a particular application will work.
03Why fix at source?
Because downstream cleaning has to run forever. A field entered inconsistently will keep being entered inconsistently, and a pipeline correcting it becomes permanent infrastructure maintaining a workaround for a process problem.
04When is cleaning the wrong answer?
When handling the defect is cheaper than removing it. A system that treats a missing field as unknown rather than requiring it to be populated may be far cheaper than backfilling three years of records nobody will read.
05How do you know when to stop?
When the use case performs acceptably against its evaluation set on real inputs. Continuing past that point is optimising data nobody is using, which is how cleanup projects turn into permanent programmes.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.