Whitepaper · 9 minute read
AI Data Platform Architecture: An Enterprise Whitepaper
An AI data platform adds four things to an analytics warehouse: curated retrieval content with ownership and freshness, a low-latency context serving layer that enforces entitlements, managed evaluation datasets, and lineage that traces any AI output back to its sources. The warehouse remains; it is not sufficient on its own.
Most enterprises approached AI with the data platform they already had, which was built for analytics: a warehouse or lakehouse holding structured data, batch pipelines, and a semantic layer for business intelligence. That platform is necessary and insufficient. AI systems need unstructured content curated for retrieval, context assembled per request within interaction latency, entitlements enforced at serving time rather than in application code, labelled data for evaluation, and lineage that traces an answer back to its sources. This whitepaper sets out what the AI data platform contains and how to build it without discarding what exists. It draws on FISTA Solutions' AI enablement practice and complements the data readiness for generative AI whitepaper and the enterprise RAG reference architecture whitepaper.
What does an AI data platform contain?
| Layer | Purpose | Not provided by the warehouse |
|---|---|---|
| Source integration | Structured and unstructured ingestion | Unstructured content at scale |
| Retrieval content store | Curated, owned, versioned content with indexes | Content curation and authority |
| Context serving | Per-request assembly with entitlements | Low-latency entitled assembly |
| Feature and state serving | Current values for decisions and agents | Interaction-latency reads |
| Evaluation data | Reference sets, labels, production samples | Managed evaluation assets |
| Lineage and provenance | Output to source traceability | Traceability through AI steps |
| Governance | Classification, entitlements, retention | Entitlement enforcement at retrieval |
| Observability | Content freshness, retrieval quality, drift | AI-specific data quality signals |
Why does the warehouse fall short?
Three mismatches. Latency: analytical queries are measured in seconds and agents need context in tens of milliseconds because they may make several calls per interaction. Data type: warehouses handle structured data well and unstructured content poorly, while most retrieval value sits in documents. And access model: warehouse security is typically role-based at table or column level, while AI serving needs per-request entitlement filtering against the end user's specific permissions, evaluated before retrieval so ranking never touches unauthorised material.
None of that makes the warehouse obsolete. It remains the source of truth for structured data, the place analytics happens, and frequently the upstream source for features and records the AI platform serves. The AI platform sits alongside and reads from it.
What makes retrieval content a managed asset?
Ownership, authority, and freshness, none of which come from indexing a repository.
An explicit inclusion model is the core discipline: content enters the retrievable corpus because a named owner designated it authoritative for a purpose, not because it happened to sit in an indexed location. That single decision prevents the most common quality failure, which is retrieval surfacing drafts, superseded versions, and personal notes with the same confidence as current policy.
Around it sit the supporting mechanics: freshness expectations per source with review cadences, authority marking so the current version is identifiable, version history retained but excluded from retrieval, and monitoring that flags content past its review date. See the enterprise knowledge management whitepaper.
What does the context serving layer do?
Assembles what the model sees, per request, within latency budget, with entitlements applied. Concretely: resolve the requesting identity and its entitlements; scope retrieval to what that identity may see; retrieve and rank; fetch the structured record fields the task needs; read current state; apply budgets and compression; render deterministically; and emit a trace of everything included with its provenance.
Two properties matter most. Entitlement filtering before retrieval, not after, because post-filtering leaks through ranking and result counts. And determinism, so the same inputs produce the same context, which makes behaviour reproducible and explicable. See the context engineering whitepaper.
Why are evaluation datasets a platform concern?
Because they gate production. A reference set determines whether a change ships and whether a regression is caught, which makes it as consequential as any production dataset and deserving of the same treatment: versioning so results are comparable over time, ownership so it is maintained, provenance for each case, access control where it contains real customer data, and refresh as the domain changes.
Teams that keep evaluation sets in spreadsheets on individual laptops discover the consequences when the person leaves, when results cannot be reproduced, or when a set containing customer data turns out to have been shared widely. The platform should hold them, version them, and serve them to the evaluation harness.
What lineage is required?
A complete path from output to source. For any AI response: which context was assembled, which documents and records it contained, which versions of those, which pipelines produced them, and from which source systems. For any action an agent took: the same, plus the instruction chain and the human whose request initiated it.
This serves three purposes that each justify it independently. Debugging, because diagnosing a wrong answer without knowing what the model saw is guesswork. Audit, because regulated organisations must explain outputs. And incident response, because determining the blast radius of a bad document or a poisoned record requires knowing what consumed it.
Conventional data lineage tooling covers the pipeline portion and stops at the platform boundary. The AI-specific extension is the context assembly step, which must be logged with document and version identifiers at request time. See what is data lineage in ai.
How are entitlements modelled?
As a first-class platform concern rather than application logic. Each content item and record carries the access control information needed to decide who may see it, synchronised from source systems and refreshed when permissions change. The serving layer evaluates the requesting identity against that before retrieval.
The hard parts are synchronisation latency, since a permission revoked in the source system must take effect in retrieval quickly, and derived artifacts, since embeddings and summaries inherit the access constraints of their sources and must be removed or re-scoped when those change.
Systems that snapshot permissions at ingestion and never refresh them are a disclosure incident with a delay, and the delay is usually measured in months.
What observability does the data layer need?
Signals specific to AI consumption, which conventional data quality monitoring does not produce: content freshness against review cadence; retrieval score distributions, which drift before answer quality visibly falls; coverage, meaning the proportion of asked questions the corpus can answer; citation concentration, which indicates whether the corpus is being used broadly or narrowly; entitlement sync lag; and context assembly latency at the tail.
These are leading indicators. A platform that monitors them catches degradation before users report it; a platform that monitors only pipeline success finds out from complaints.
How should this be built incrementally?
Not as a platform programme. The sequence that works starts with the first use case and extracts shared components as the second and third arrive.
- Content curation for one domain, with owners and authority marking, serving the first retrieval use case.
- Context serving extracted as a component when the second use case appears, with entitlement filtering built in.
- Evaluation data management once more than one reference set exists.
- Lineage logging at context assembly from the start, since retrofitting it is painful.
- Entitlement synchronisation hardened as the content scope broadens beyond what one permission model covers.
- Observability added as production volume makes degradation consequential.
Organisations that attempt the full platform before any use case build for imagined requirements and deliver late; organisations that never extract shared components rebuild the same serving logic per project with inconsistent security.
How does this relate to the data mesh or fabric debate?
Largely orthogonally. Whether content and data are owned centrally or by domains, the AI platform still needs curation with ownership, entitled serving, evaluation management, and lineage. Domain ownership arguably suits retrieval content better than central ownership, because the people who know whether a policy is current are in the domain.
What does not work is treating retrieval content as an unowned by-product of domain systems, which is the arrangement most organisations start with and the reason their first retrieval system degrades.
What goes wrong?
Indexing repositories rather than curating content. Entitlements filtered after retrieval or snapshotted at ingestion. Evaluation sets on laptops. No lineage through context assembly, so wrong answers cannot be diagnosed. Latency requirements discovered after the warehouse was chosen as the serving layer. And platform programmes that run for a year before a use case reaches production.
Who owns the AI data platform?
Usually the existing data platform team, extended rather than duplicated, because the alternative is two organisations with overlapping responsibilities arguing about which owns the pipeline that feeds retrieval. What the team needs to add is competence in unstructured content handling, entitlement modelling at request granularity, and the latency discipline that serving agents requires, none of which analytics engineering develops naturally.
Content ownership sits elsewhere, with the domains. The platform team provides the mechanism, freshness monitoring, authority marking, ingestion, and the domains provide the judgement about what is authoritative and current. Platforms that try to own content accuracy centrally fail, because a central team cannot know whether a regional policy was superseded last week.
The role that most often goes unfilled, and that determines whether the platform stays healthy, is an owner for the evaluation datasets. They are nobody's obvious responsibility, they decay quietly, and their decay is invisible until a regression ships.
What does the first year look like?
Quarter one: content curation for the first domain with named owners, ingestion and indexing, and lineage logging built in from the first request. Quarter two: the first retrieval use case in production, with entitlement filtering enforced at serving and freshness monitoring running. Quarter three: context serving extracted as a shared component as the second use case arrives, plus evaluation dataset management moved off individual machines. Quarter four: entitlement synchronisation hardened as scope widens, observability on retrieval score distributions and coverage, and a review of which shared components the next wave of use cases will need.
The test at the end of the year is whether the second use case cost materially less than the first. If it did not, the shared components were not extracted and the organisation is building projects rather than a platform.
How FISTA Solutions delivers this
FISTA Solutions builds AI data platforms incrementally alongside existing warehouses, with curated owned retrieval content, entitlement-filtered context serving, managed evaluation datasets, and lineage from output to source, so AI systems are explicable and governable as they scale, through AI enablement, AI agents, and forward deployed engineers working with data teams. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To build the data foundation AI actually needs, message FISTA on WhatsApp, or read the data readiness for generative AI whitepaper.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01Why is an analytics warehouse insufficient for AI?
Because it is optimised for batch analytical queries over structured data, while AI systems need low-latency retrieval of unstructured content, entitlement-filtered context assembly per request, and reference data for evaluation, none of which a warehouse is designed to provide.
02What is a context serving layer?
The component that assembles what a model sees for a given request: retrieved passages, structured record fields, and current state, fetched at interaction latency with the requesting user's entitlements applied before retrieval rather than after.
03How should retrieval content be managed?
As a curated asset with named owners per domain, freshness expectations, authority marking so superseded material is excluded, and an explicit inclusion model where content enters because someone designated it rather than because a repository was indexed.
04Why do evaluation datasets need governance?
Because they determine whether systems ship and whether regressions are caught, which makes them production assets. They need versioning, ownership, provenance for each case, access control where they contain real customer data, and refresh as the domain changes.
05What lineage does AI require?
A path from any output back through the assembled context to the source documents and records, and from those to the systems and pipelines that produced them, sufficient to answer what an answer was based on and whether the requester was entitled to it.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.