FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Glossary · 5 minute read

What Is Purpose Limitation? Using Data Only as Intended

Purpose limitation requires that personal data collected for a stated purpose is not further processed in a way incompatible with it. AI projects collide with this constantly, because they reuse data collected for service delivery to train or ground systems built for different purposes.

By FISTA Solutions· AI-Native Engineering Team·
What Is Purpose Limitation? Using Data Only as Intended article cover

Purpose limitation is the principle most likely to stop an AI project, and it is usually encountered after the project is underway. The reason is structural: AI projects succeed by reusing data an organisation already holds, and that data was collected for something else. This explainer covers how the principle applies. It complements what is a data protection impact assessment and ai and gdpr, and reflects FISTA Solutions' approach in AI enablement delivery. This article is general guidance, not legal advice.

What does the principle require?

That personal data collected for a specified purpose is not further processed in a manner incompatible with that purpose. Further processing is permitted where it is compatible, and compatibility is assessed rather than asserted.

The assessment weighs the relationship between purposes, the context in which the data was collected, the nature of the data, the possible consequences for individuals, and the safeguards in place.

UseTypical compatibilityNote
Answering a user's question about their own dataUsually compatibleClose to original purpose
Internal analytics on aggregatesOften compatibleDepends on safeguards
Training a model on customer recordsOften a new purposePersists in weights
Building a product feature for other customersUsually newDifferent beneficiary
Sharing with a third partyNew purposeRequires its own basis

Why is training usually a new purpose?

Because the data was collected to deliver a service, not to build an artefact. The consequences differ materially: training embeds the data in a model that persists, may be difficult to unwind, and may be used in contexts unrelated to the original relationship.

Treating training as a natural extension of service delivery is a position organisations take and rarely defend well when examined. Considering it a new purpose, and establishing a basis for it, is the more sustainable stance.

Does broad consent cover it?

Usually less than hoped. Purposes stated broadly enough to cover anything tend to be treated as insufficiently specific, and consent obtained years ago for service delivery is a weak foundation for a model built afterwards.

This matters practically because many organisations discover their strongest-looking basis is a privacy notice that predates the technology entirely.

Why is retrieval often easier?

Because it resembles the original purpose more closely. Retrieving a document to answer a user's question about it is close to what the data was collected for, and the processing is transient, scoped to that user, and subject to access control.

Training the same document into a model is a different operation with different persistence and different reach. Where the requirement is knowledge rather than behaviour, the retrieval architecture is frequently both better engineering and a more defensible position. See fine-tuning vs rag.

What about anonymising first?

It helps if genuinely achieved, and genuine anonymisation is a high bar that most pipelines do not reach. Pseudonymised data remains personal data, so the principle still applies, and de-identified free text frequently retains identifying detail.

Anonymisation is worth pursuing and is not a shortcut around establishing the basis.

When should this be settled?

Before the project begins. Establishing the basis afterwards means building on an unresolved question, and discovering late that the data cannot be used wastes the whole effort.

The conversation is also more productive early, because the design can adapt — a different data source, a retrieval architecture, an aggregate rather than record-level approach — in ways that are impossible once the system is built.

What should you do first?

For your highest-value AI use case, write down what the data was originally collected for and what you now intend to do with it. If those two sentences describe different things, you have a purpose question to resolve, and resolving it now is cheaper than resolving it later.

How does this apply to feedback data?

Interaction logs are the data organisations most want to reuse and most often reuse without thinking. Conversations with a support assistant were generated to resolve an issue; using them to train a model is a different purpose, and the content frequently includes personal information the user volunteered.

Using them for evaluation is usually easier to defend than using them for training, because evaluation is transient, reviewable, and does not embed anything permanently. Distinguishing the two in policy, rather than treating all secondary use identically, gives teams a workable path.

What about vendor terms?

They matter in both directions. A provider's terms determine whether your data may be used to improve their models, which is a purpose your customers did not agree to and which enterprise agreements usually exclude. Checking that exclusion, per model and per tier, is part of establishing your own position rather than an administrative detail.

How FISTA Solutions helps

FISTA Solutions establishes the lawful basis and purpose position before design, prefers retrieval architectures where they are both better engineering and more defensible, avoids relying on broad historical consent language, and adapts data sourcing where the purpose question cannot be resolved, through AI enablement, AI agents, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies across 12+ countries.

To resolve data reuse questions before they stop your project, message FISTA on WhatsApp, or read what is a data protection impact assessment.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01What does compatibility mean?

An assessment of the relationship between the original and new purposes, the context of collection, the nature of the data, possible consequences for individuals, and safeguards applied. It is a judgement, not a checkbox, and it is where most AI projects need a considered answer.

02Is training a new purpose?

Usually yes. Data collected to deliver a service was not collected to build a model, and the two differ in consequence: training embeds the data in an artefact that persists and may be difficult to unwind. Treating them as the same purpose is generally optimistic.

03Does broad consent language help?

Rarely as much as hoped. Purposes stated broadly enough to cover anything are often treated as insufficiently specific, and consent given years ago for service delivery is a weak basis for a model built afterwards. This is general guidance, not legal advice.

04Is retrieval easier than training?

Frequently, yes. Retrieving a document to answer a user's question about it resembles the original service purpose far more closely than embedding it permanently in model weights, and it preserves deletion, access control, and the ability to withdraw a document from scope.

05When should this be settled?

Before the project starts. Establishing the basis afterwards means either building on an unresolved question or discovering late that the data cannot be used, and both outcomes are considerably more expensive than having the conversation at the design stage.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project