Cost · 5 minute read
AI Data Cleanup Cost: The Work Before Any Model Runs
Data cleanup is typically the largest cost in an AI project and the least visible in its plan. Entity resolution, definition reconciliation across systems, and handling historical inconsistency all require judgement that cannot be automated away, and scoping them requires knowing what the data will be used for.
Data cleanup is the largest cost in most AI projects and appears in almost none of their plans. Organisational data was collected to run operations, not to train models or answer analytical questions, and making it coherent enough to rely on is substantial work requiring judgement that no tool supplies. This guide covers what that work involves, drawing on FISTA Solutions' AI enablement delivery. It complements ai data readiness and how to build a data quality agent.
Why does cleanup dominate?
Because the data was collected for something else. Operational systems record what they need to operate, with conventions that suited the moment, fields repurposed when requirements changed, and no obligation to be consistent with any other system.
Making that coherent enough to rely on is the bulk of most AI projects, and it is invisible in plans that scope modelling and integration.
| Task | Effort | Automatable |
|---|---|---|
| Entity resolution | Large | Partly, with review |
| Definition reconciliation | Large | No, requires decisions |
| Historical inconsistency handling | Large | No, requires decisions |
| Format and type normalisation | Moderate | Largely |
| Deduplication | Moderate | Partly |
| Gap identification and filling | Variable | Identification yes, filling no |
What is entity resolution?
Determining that records in different systems refer to the same customer, supplier, product, or asset. It is the recurring hard problem in enterprise data.
Identifiers differ between systems, names vary in spelling and format, addresses change, and duplicates exist within systems as well as across them. Automated matching handles the clear cases and produces a set of ambiguous ones that require human judgement — and that review is the cost.
Why do definitions differ?
Because each system was built for its own purpose by people solving their own problem. Active customer means one thing in the billing system and another in the CRM. Revenue is recognised differently. A completed order means dispatched in one system and delivered in another.
Aggregating across those without reconciling produces numbers nobody can defend, and the reconciliation is a business conversation rather than a technical task.
What makes historical inconsistency hard?
That resolving it requires decisions rather than code. When a field's meaning changed three years ago, or a product hierarchy was reorganised, someone must decide how to treat the earlier records — restate them, exclude them, or carry the discontinuity explicitly.
Each choice affects every downstream number, and the decision belongs to whoever is accountable for those numbers. Projects that treat it as a technical problem make the decision implicitly and discover it when a figure is challenged.
How should cleanup be scoped?
By the use case. Cleaning everything is an open-ended programme that never completes; cleaning what a specific use case requires is a bounded project with a clear completion criterion.
That means specifying the use case first and deriving the data requirement from it, rather than beginning a data quality initiative and hoping it enables something. Fitness for the stated purpose is the standard. See ai data readiness.
What can automation contribute?
Detection and proposal rather than resolution. Identifying probable duplicates, flagging definition mismatches, surfacing values outside expected ranges, and proposing entity matches all narrow what humans must examine.
That narrowing is genuinely valuable and it does not remove the judgement. A matching system that resolves ambiguous cases automatically produces a clean-looking dataset with wrong merges in it, which is harder to detect than the original mess.
What is the ongoing cost?
Continuous, because data keeps arriving. Cleanup that treats the existing corpus and does not address the pipelines producing new data delivers a clean snapshot that degrades from the day it is finished.
Addressing quality at source — validation at entry, consistent definitions in new systems — is what makes the cleanup last.
What should you do first?
Take the specific question your AI use case must answer and trace the data it needs back to source. That exercise identifies exactly which cleanup is required and, frequently, that a narrower use case is achievable much sooner.
Who should do the work?
People who know what the data means, supported by engineers who can move it. Cleanup performed entirely by a technical team produces a dataset that is internally consistent and subtly wrong, because the decisions about what a field means and how to treat a discontinuity were made by people without the context to make them.
That requirement is what makes cleanup hard to resource: it needs time from operational staff who have other jobs, and their availability rather than engineering capacity is usually the constraint.
How does this affect project estimates?
Substantially, and in a direction estimates rarely go. A project scoped at three months of engineering frequently needs a further two of data work that nobody surveyed, and the discovery happens after commitments have been made.
Spending a week assessing data state before committing to a schedule is the cheapest risk reduction available in AI project planning, and it is skipped almost universally because it delays the start of visible work.
How FISTA Solutions helps
FISTA Solutions scopes data cleanup from the use case rather than as an open-ended initiative, uses automation to narrow what requires human judgement without resolving ambiguity automatically, treats definition reconciliation as a business decision, and addresses source pipelines so cleanup lasts, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 47% efficiency gains.
To scope the data work before it scopes your project, message FISTA on WhatsApp, or read ai data readiness.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why does cleanup dominate?
Because organisational data was collected for operational purposes rather than for analysis or training. It is inconsistent across systems and across time, and making it coherent enough to rely on is substantial work that precedes anything else.
02What is entity resolution?
Determining that records in different systems refer to the same customer, supplier, or product. It is the recurring hard problem because identifiers differ, names vary, and duplicates exist within systems as well as between them.
03Why do definitions differ?
Because each system was built for its own purpose. Active customer, revenue, and completed order all mean subtly different things in different systems, and aggregating them without reconciling produces numbers nobody can defend.
04What makes historical inconsistency hard?
That fixing it requires business decisions rather than code. When a field's meaning changed three years ago, someone must decide how to treat the earlier records, and that decision affects every downstream number.
05How should cleanup be scoped?
By the use case. Cleaning everything is an open-ended programme; cleaning what a specific use case requires is a bounded project with a clear completion criterion. Fitness for purpose is the standard, not abstract quality.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.