Checklist ┬╖ 5 minute read
AI Data Migration Checklist: Moving a Corpus Without Loss
Corpus migrations lose content quietly. Inventory what exists before moving anything, map permissions explicitly, plan re-embedding as a cost and a schedule, validate by comparing retrieval results before and after, and keep the old system available until the new one has been verified under real traffic.
Corpus migrations lose content quietly, and the loss is usually discovered by a user rather than by the team. This checklist covers doing it verifiably, drawn from FISTA Solutions' AI enablement delivery work.
What does a migration involve?
Six phases, each with a verification step.
| Phase | Verification |
|---|---|
| Inventory | Counts and checksums recorded |
| Permission mapping | Access tested per role |
| Transfer | Every item accounted for |
| Re-embedding | Coverage confirmed, cost tracked |
| Validation | Retrieval compared side by side |
| Cutover | Rollback available and tested |
Inventory and baseline
You cannot verify a migration you did not measure first.
- Document count, chunk count, and total size recorded
- Source systems and their owners listed
- Metadata schema documented for every field in use
- A checksum or identifier recorded per document
- Known-broken or orphaned content identified before the move
- A query set assembled for before-and-after comparison
- Baseline retrieval results captured for that query set
Permissions and access
This is where migrations cause security incidents. Map explicitly rather than assuming equivalence.
- Permission model of both systems documented
- Mapping written for every role and access level
- Documents with restricted access identified explicitly
- Access tested per role after transfer, not assumed
- Over-granting checked as carefully as under-granting
- Service account permissions scoped to the migration only
- Permission changes logged for audit
Transfer and integrity
Every item must be accounted for, including the ones that failed.
- Transfer is resumable from any interruption point
- Every document accounted for: transferred, skipped, or failed
- Failures logged with reasons, not silently dropped
- Metadata preserved including dates, sources, and ownership
- Character encoding verified on a sample of non-English content
- Large or unusual documents tested explicitly
- Counts reconciled against the inventory baseline
Re-embedding
A full pass over the corpus with real cost and real elapsed time.
- Whether re-embedding is required determined and stated
- Compute cost estimated before starting
- Elapsed time estimated and scheduled
- Progress tracked so partial completion is visible
- Embedding model version recorded with the index
- Dimension and index configuration verified against the new model
- A sample of embeddings spot-checked for sanity
Validation
Compare behaviour, not counts. See RAG quality checklist.
- The baseline query set run against the new system
- Retrieved results compared item by item against the baseline
- Differences investigated rather than accepted
- Retrieval latency measured and compared
- Evaluation suite re-run end to end
- A domain expert reviews a sample of answers from both systems
- Permission-filtered retrieval tested for each role
Cutover and rollback
Keep the reverse path available until the forward path is proven.
- Cutover plan written with a rollback trigger
- Old system kept available and readable after cutover
- Traffic moved gradually where possible rather than all at once
- Monitoring in place for retrieval failures and latency
- A defined watch period before decommissioning anything
- Users told what is changing and how to report problems
- Decommissioning scheduled explicitly rather than forgotten
What are the most common failures?
Verifying by document count. Assuming permissions map cleanly. Discovering re-embedding cost mid-migration. Decommissioning the old system immediately. And no baseline query set, which makes validation impossible.
Who should own this?
Engineering runs the migration; the content owner validates that the corpus still answers what it should. Without the second, the migration is verified technically and not functionally.
How often should it run?
Per migration. The baseline query set and validation approach should be reusable, because migrations recur тАФ provider changes, embedding model upgrades, and consolidation all trigger one.
What evidence should it produce?
Inventory counts before and after, the permission mapping, transfer logs including failures, and the side-by-side retrieval comparison. That package demonstrates nothing was lost.
What if you find loss after cutover?
Roll back if the old system is still available, which is the argument for keeping it. If it is not, re-migrate the missing content and treat the gap as an incident.
The more common situation is discovering loss weeks later through user reports. Continuous comparison of retrieval results for a period after cutover is what catches it earlier. See observability for web apps.
What should you do first?
Capture baseline retrieval results for fifty representative queries before you move anything. That single artefact is what makes validation possible.
How FISTA Solutions helps
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: migrations validated by comparing retrieval results against a captured baseline rather than by counting documents, with the old system kept readable through a watch period, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To adapt this checklist to your environment, message FISTA on WhatsApp, or read RAG quality checklist.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What goes wrong most often?
Silent partial loss. Documents fail to migrate, chunks are dropped, or permissions do not carry over, and nothing errors. The system works and quietly cannot answer questions it used to.
02Why are permissions difficult?
Because permission models differ between systems. A mapping that looks equivalent often grants slightly more or slightly less, and both directions cause problems that surface later.
03When is re-embedding required?
Whenever the embedding model changes. It is a full pass over the corpus with real compute cost and real elapsed time, and it must be planned rather than discovered mid-migration.
04How do you validate properly?
Run the same query set against both systems and compare retrieved results. Matching document counts proves nothing about whether retrieval still works.
05How long should the old system stay?
Until the new one has handled real traffic for long enough to expose problems тАФ usually a few weeks. Decommissioning early removes the only fast rollback available.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.