Checklist ¡ 4 minute read
AI Data Map Template: Tracing Data Through the System
A data map answers where data came from, where it went, who saw it, and on what basis. For AI systems that means tracing inputs through retrieval, model providers, logs, and derived stores such as embeddings â which is where most maps stop short.
A data map answers where data came from, where it went, and who saw it. This template covers building one for AI systems, drawn from FISTA Solutions' AI enablement governance work. This is general guidance, not legal advice.
What does the map cover?
Six stages, each with its own record.
| Stage | What to record |
|---|---|
| Collection | Source, basis, categories |
| Processing | What happens, by whom |
| Retrieval and context | What is pulled and shown to the model |
| Third-party flows | Providers, locations, terms |
| Storage | Every store including derived ones |
| Retention and deletion | Period and mechanism per store |
Sources and collection
Where data enters the system.
- Every input channel identified
- Data categories per channel recorded
- Lawful basis per channel recorded
- Whether the person knows this data is processed by AI
- Notice or consent reference per channel
- Special category data identified
- Data obtained from third parties recorded separately
Processing and flows
What happens to it and where it goes.
- Processing steps listed in order
- What is sent to the model provider, precisely
- Retrieved content included as a flow
- Tool calls that send data externally recorded
- Actions written back to other systems recorded
- Any human review step recorded
- Data minimisation applied and documented
Third parties
Every party that sees the data. See AI subprocessor checklist.
- Model providers listed with what they receive
- Inference hosts listed
- Observability and logging tools that see content listed
- Storage providers listed
- Human review services listed
- Processing locations per party
- Transfer mechanisms recorded
Stores
Including derived ones, which is where maps usually stop.
- Primary application store
- Search index
- Vector store and embeddings
- Prompt and response logs
- Caches
- Analytics and warehouse copies
- Backups and exports
Retention and deletion
Per store, with the mechanism. See AI data retention checklist.
- Retention period recorded per store
- Deletion mechanism recorded per store
- Stores unable to delete selectively flagged
- Deletion request path traced through the map
- Third-party deletion obligations recorded
- Backup handling documented
- Verification approach recorded
Maintenance
Tied to change, not to audit season.
- Map owner named
- Update required when architecture changes
- Linked to the service catalog entry
- Reviewed at least annually
- Reviewed when a provider or subprocessor changes
- Version history kept
- Accessible to privacy, security, and engineering
What are the most common failures?
Mapping the primary store only. Treating model providers as infrastructure. Omitting logs. Basis recorded once rather than per step. And a map produced for an audit and then abandoned.
Who should own this?
A privacy or governance function owns the map's format; the system's technical owner owns its accuracy. Maps maintained without engineering input describe an architecture that no longer exists.
How often should it run?
Created before launch, updated on architecture change, and reviewed annually. Provider changes should trigger an update to the third-party section.
What evidence should it produce?
The current map with a review date, its linkage to the service catalog, and a traced deletion request demonstrating the map is accurate.
How do you verify the map is correct?
Trace a real record through it. Pick an identifier, find it in every store the map lists, and check whether it appears anywhere the map does not.
That exercise routinely finds a store nobody documented, which is exactly what the map exists to prevent. Do it before an auditor does. See AI data retention checklist.
What should you do first?
List every store holding data from one AI system, including embeddings and logs. The list is usually longer than the map you already have.
How FISTA Solutions helps
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: data traced through derived stores including embeddings and logs, with lawful basis recorded per processing step rather than once at collection, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To adapt this checklist to your environment, message FISTA on WhatsApp, or read AI subprocessor checklist.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What does the map need to show?
Every point data enters, every store it rests in including derived ones, every third party it reaches, the basis for each processing step, and the retention applied at each point.
02Where do maps usually stop short?
At derived stores. Embeddings, caches, search indexes, and logs all hold data traceable to individuals, and they are routinely omitted from maps that cover the primary database carefully.
03Why are logs significant?
Because prompt and response logs hold the input, the retrieved content, and the output together â frequently the most complete and sensitive copy in the system.
04Why record basis per step?
Because it can differ. Collection, processing by a model provider, retention for audit, and use for improvement may each rest on a different basis, and the map is where that is visible.
05How is it kept current?
By tying updates to architecture change and to the catalog entry, and by reviewing at least annually. A map maintained only during audits is stale when it matters.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.