Glossary · 4 minute read
What Is Prompt Leaking? System Prompt Exposure Explained
Prompt leaking is the extraction of a system prompt through user interaction, exposing instructions the operator assumed were private. It cannot be reliably prevented, which means system prompts must never contain secrets, credentials, or access rules, and security must be enforced outside the prompt entirely.
Prompt leaking is discussed as a vulnerability and is better understood as a design assumption. System prompts will eventually be extracted from any application with enough users and enough interest, and the question worth answering is not how to prevent it but what breaks when it happens. This explainer covers both. It complements what is prompt injection and ai security checklist, and reflects FISTA Solutions' approach in AI agents delivery.
How does extraction work?
By reframing. Rather than asking for the prompt directly, which most applications now block, a user asks for a summary of the instructions given, a translation of the preceding text, a continuation of what came before their message, or engages in a role-play where discussing the configuration is in character.
New phrasings appear faster than filters can be written, which is the structural reason instruction-based prevention does not hold.
| Assumption | Safe | Why |
|---|---|---|
| Prompt stays private | No | Extraction techniques evolve |
| Secrets in prompt are hidden | No | Extracted with the prompt |
| Prompt enforces access control | No | Instruction, not enforcement |
| Prompt is product design | Yes | Treat as competitively sensitive |
| Prompt is public | Yes | The correct default assumption |
Can it be prevented?
Not reliably. Instructions to keep the prompt confidential raise the effort required and do not close the category. Output filters that detect prompt text can be circumvented by asking for it transformed — translated, paraphrased, encoded, or summarised.
Defence in depth is worth having, and treating it as sufficient is the error. The design principle is that exposure should be embarrassing rather than damaging.
What must never be in a prompt?
Credentials of any kind. Internal endpoints and infrastructure details. Database schemas and connection information. Business rules whose disclosure would enable circumvention. Personal data. Anything an attacker would find useful.
The test is simple: if publishing this text would be a problem, it does not belong in a prompt. That test is failed routinely, most often with internal URLs and identifiers included for the model's convenience.
Why is prompt-based access control invalid?
Because an instruction is a request. "Only answer questions about the user's own account" describes desired behaviour; it does not prevent the model from answering otherwise when the context makes that seem appropriate.
Real enforcement scopes retrieval to the authenticated identity before search, checks authorisation in the tool implementation, and validates output against rules in code. Those hold regardless of what the model is persuaded to attempt. See ai access control.
What does leakage actually cost?
In a correctly built system, competitive advantage. A good prompt represents real product work — the phrasing, the examples, the edge cases handled — and exposing it hands that to competitors.
In an incorrectly built system, considerably more: exposed credentials, disclosed internal architecture, and a published description of the rules a user would need to circumvent. The difference is entirely in what was put in the prompt.
How does this relate to prompt injection?
They are related and distinct. Leaking extracts the prompt; injection overrides it, usually through content the system retrieves rather than through direct user input. Both defeat the assumption that instructions in the context are authoritative.
The shared lesson is that the context window is not a trust boundary. Anything that must hold has to be enforced outside it.
What should teams actually do?
Keep prompts free of anything sensitive, enforce security in code, and accept that the prompt may become public. Some teams go further and publish their system prompts deliberately, which removes the concern entirely and signals confidence in the rest of the design.
What should you do first?
Read your production system prompts looking for anything you would not want published — URLs, identifiers, rules, credentials, data. Most teams find at least one item. Removing those converts prompt leaking from a security issue into a competitive annoyance, which is where it should sit.
Does model choice affect this?
Marginally. Models differ in how readily they disclose their instructions, and providers continue to improve resistance, but none of them offers a guarantee and none should be relied on for one. Choosing a model for its resistance to extraction is optimising a defence that was never meant to be load-bearing.
What does vary meaningfully is behaviour under injection, where a model's tendency to follow instructions found in retrieved content differs and is worth evaluating for agent systems.
How does this affect multi-tenant products?
More than in single-tenant applications, because one customer extracting a prompt may learn how other customers' configurations work. Where prompts are customised per tenant, tenant-specific detail in them becomes a cross-tenant disclosure risk, and the safe pattern keeps per-tenant configuration in data the application applies rather than in prompt text the model holds.
How FISTA Solutions helps
FISTA Solutions builds applications on the assumption that prompts are public, keeps credentials and infrastructure detail out of them entirely, enforces access control through entitlement-scoped retrieval and authorisation in tool implementations, and validates output in code rather than by instruction, through AI agents, AI enablement, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To make prompt exposure harmless, message FISTA on WhatsApp, or read ai security checklist.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01How do system prompts get extracted?
Through requests that reframe the task — asking for a summary of instructions, a translation, a continuation of preceding text, or a role-play in which the prompt is discussed. New phrasings appear faster than filters can be written, which is why prevention by instruction fails.
02Can leaking be prevented?
Not reliably. Instructions telling a model to keep its prompt secret raise the difficulty and do not close the class. Any design whose security depends on the prompt remaining hidden is resting on something that will eventually fail.
03What must never appear in a prompt?
API keys, credentials, internal endpoints, database details, unpublished business rules with security implications, personal data, and anything whose disclosure would matter. The test is simple: if publishing this text would be a problem, it does not belong in a system prompt at all.
04Why is prompt-based access control invalid?
Because an instruction is a request, not an enforcement. A prompt saying only answer questions about the user's own account is a hope; retrieval scoped by authenticated identity is a control. The difference shows the first time someone tries.
05What does exposure actually cost?
Usually competitive rather than security harm: a prompt is product design and effort, and revealing it hands that work to anyone. Where it has been built correctly, that is the whole cost, which is the position to aim for.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.