Whitepaper · 8 minute read
Building an AI Red Teaming Program: An Enterprise Whitepaper
An AI red teaming programme systematically attacks your own AI systems to find failures functional testing misses: prompt injection, tool abuse, data exfiltration, guardrail evasion, and harmful output. It needs a defined scope and threat model, a repeatable attack library, mixed internal and external exercises, triage into engineering backlogs, and a cadence tied to change rather than the calendar.
Security testing for conventional software assumes the attacker is outside the system, probing interfaces. AI systems break that assumption: the input is instructions, the retrieved document may be written by an adversary, and the agent holds credentials to act. A system that passes every functional test and scores well on evaluation can often be redirected by a sentence hidden in a document it was asked to summarise. Finding that requires deliberate adversarial work. This whitepaper sets out how to build the programme that does it. It draws on FISTA Solutions' security practice across AI agents deployments and complements what is ai red teaming and the AI agent security architecture whitepaper.
What is in scope, and what is the threat model?
Scope covers every AI system that reaches untrusted input or holds capability worth abusing: customer-facing assistants, internal agents with tool access, retrieval systems over shared content, document processing pipelines, and any system whose output feeds a decision or an action.
The threat model should name adversaries explicitly. External users probing a public assistant. Insiders escalating access through an agent that holds broader permissions than they do. Third parties who can place content into a system the agent reads, which includes anyone who can send an email, file a ticket, submit a document, or edit a wiki page. And automated adversaries testing at scale. The last two are the ones enterprise threat models routinely omit, and they are where the damage happens.
What are the attack categories?
| Category | What the attacker attempts | Consequence |
|---|---|---|
| Direct prompt injection | Override instructions through the user input | Policy bypass, unauthorised output |
| Indirect prompt injection | Hide instructions in content the agent retrieves | Agent acts for the attacker with its own permissions |
| Tool abuse | Induce calls with attacker-chosen arguments | Unauthorised reads, writes, transactions |
| Data exfiltration | Extract other users' data, secrets, or system prompts | Confidentiality breach |
| Permission escalation | Use the agent's credentials to exceed the user's own | Access control bypass |
| Guardrail evasion | Reframe, encode, or split a prohibited request | Harmful or prohibited output |
| Output-based attack | Produce content that attacks a downstream system | Injection into SQL, shells, browsers |
| Denial and cost abuse | Trigger expensive loops or unbounded generation | Availability and cost impact |
| Poisoning | Introduce content that corrupts retrieval or memory | Persistent wrong behaviour |
| Social engineering via agent | Use the agent's authority to mislead a person | Fraud, bad decisions |
The second row deserves emphasis. Indirect prompt injection is the attack that makes agents categorically different from chatbots: an attacker who can get text into a document, email, ticket, or web page that the agent reads inherits the agent's tool permissions. Defences are partial, which means containment through permissions and approval gates matters more than detection. See what is indirect prompt injection and the prompt injection defense checklist.
How is an exercise designed?
With objectives, not vibes. A useful exercise states what the team is trying to achieve, for example: cause the agent to reveal another tenant's data; cause a write action outside the user's entitlement; obtain the system prompt; cause a policy-violating output that passes the guardrails; or drive cost above a threshold with a single request.
Each objective gets a defined success criterion, a scope boundary including which environments and data may be used, and a rule of engagement covering what testers may not do, particularly where production systems or real customer data are involved. Most AI red teaming should run against a production-equivalent environment with synthetic data, escalating to production only with explicit authorisation and monitoring.
What does the attack library contain?
A maintained, versioned set of techniques with expected system behaviour, so that exercises are repeatable and regressions are detectable. It grows from three sources: published research and disclosed techniques, findings from the organisation's own exercises, and incidents.
The library's real value is regression. Every confirmed finding becomes a permanent test case that runs on every model change, prompt change, and tool addition, which is what prevents the same vulnerability reappearing after a refactor. Without this, red teaming is a series of reports; with it, it is a control.
Who runs the exercises?
A mix, deliberately. Internal engineers know the architecture and find system-specific weaknesses fastest, but they share the design assumptions that created the vulnerabilities. External specialists bring current technique and no assumptions, at higher cost and with less system context. Domain experts, meaning the people who understand the business consequences, identify harms that technically minded testers miss entirely, such as outputs that are technically permitted and professionally catastrophic.
A practical structure runs internal exercises continuously as part of engineering, external assessments periodically and before major launches, and domain expert review sessions on the outputs that matter most.
How are findings managed?
Like security findings, because they are. Each gets a severity based on exploitability and impact, an owner, a remediation path, and a due date, tracked in the engineering backlog rather than a report. Severity for AI findings should weight blast radius heavily: an injection that produces a rude response is low; one that triggers a refund or reads another tenant's record is critical regardless of how difficult it was to achieve.
Three outcomes are legitimate: fix, which usually means tightening permissions or adding validation rather than patching a prompt; mitigate, meaning containment through approval gates or monitoring; and accept with documented rationale and monitoring. What is not legitimate is a finding that closes because the prompt was edited and the specific phrasing no longer works, since the class of attack remains.
What defences actually hold?
Prompt-level defences are helpful and insufficient, because the attack surface is natural language and the defender cannot enumerate it. The defences that hold are architectural:
- Least privilege on tools, so the worst an injected instruction can achieve is bounded by what the agent may do at all.
- Approval gates on consequential actions, so a person confirms anything irreversible or material.
- Separation of untrusted content from instructions, with retrieved material clearly delimited and treated as data.
- Output validation against schemas and business rules before any action.
- Entitlement filtering before retrieval, so an agent cannot surface what the user may not see.
- Egress control, limiting where an agent can send data.
- Rate and cost limits per workflow.
- Full tracing, so any incident can be reconstructed.
Architecture is the defence; prompts are a mitigation. See ai agent guardrails and how to design tool permissions for ai agents.
What cadence does the programme run on?
Change-driven rather than calendar-driven, because AI systems change continuously. Triggers that should mandate an exercise: before any launch; when a tool or permission is added, which changes blast radius; when a new data source enters retrieval, which changes who can inject; when the model version changes, since defences are model-dependent; after any incident; and when a significant new technique is published.
A periodic full exercise, annually or semi-annually, acts as a backstop and satisfies audit expectations, but a programme that runs only on that schedule will be months behind its own system.
How does this integrate with existing security?
It belongs inside the security programme, not beside it. AI findings enter the same vulnerability management process, the same severity scale, and the same reporting. Penetration testing scope expands to include AI components. The secure development lifecycle gains AI-specific gates. Incident response playbooks gain AI scenarios.
The one genuine difference is that AI security findings frequently require product and domain judgement to assess, because whether an output is harmful depends on context that a security engineer may not hold. Building that review into triage prevents both over- and under-reaction. See ai security checklist and how to secure an ai system.
How is programme maturity measured?
Coverage, meaning the share of AI systems exercised and the attack categories tested per system. Finding rate and severity distribution over time, which should show severe findings falling as architecture improves. Time to remediate by severity. Regression suite size and pass rate. And the proportion of findings discovered internally rather than by customers or researchers, which is the outcome measure that matters.
How does a programme start?
- Inventory AI systems with their permissions, data access, and untrusted input paths.
- Rank by blast radius, not by visibility; the internal agent with write access usually outranks the public chatbot.
- Run a first exercise on the highest-risk system with three or four concrete objectives.
- Triage into the backlog with severity and owners; fix the architecture, not the phrasing.
- Build the regression suite from confirmed findings and run it in the pipeline.
- Add triggers so tool, data source, and model changes require an exercise.
- Bring in external testing before the next significant launch.
What goes wrong?
Testing only the public chatbot while internal agents hold write access to finance systems. Treating prompt hardening as the fix. Findings that live in a PDF. Exercises run once at launch on a system that has changed twenty times since. No regression suite, so fixes silently regress. And a threat model that omits the person who can file a ticket, which is the easiest way into most enterprise agents.
How FISTA Solutions delivers this
FISTA Solutions builds red teaming into AI delivery rather than bolting it on, with architectural containment through least-privilege tools and approval gates, maintained attack libraries that become regression suites, and exercises triggered by change, through AI enablement, AI agents, and forward deployed engineers working with security teams. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.
To find your AI failures before someone else does, message FISTA on WhatsApp, or read the AI agent security architecture whitepaper.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What is AI red teaming?
Structured adversarial testing of AI systems to find failures that normal evaluation misses, including prompt injection, tool misuse, data exfiltration, guardrail evasion, and harmful or policy-violating output. It tests what a motivated adversary can make the system do, not how well it performs on average.
02How is it different from evaluation?
Evaluation measures typical quality on representative inputs. Red teaming measures worst-case behaviour under adversarial inputs designed to break the system. A system can score highly on evaluation and be trivially exploitable, which is the normal state of untested agent deployments.
03What is the most important attack class for enterprises?
Indirect prompt injection, where instructions hidden in retrieved documents, emails, tickets, or web content redirect an agent that has tool access. It is the defining enterprise risk because the attacker never touches the interface and the agent's permissions become theirs.
04Who should run the exercises?
A mix. Internal engineers find system-specific weaknesses fastest because they know the architecture; external specialists bring current techniques and fresh assumptions; and domain experts identify harms engineers miss. Purely internal programmes converge on their own blind spots.
05How often should red teaming run?
On change rather than on a schedule: before launch, when new tools or permissions are added, when new data sources enter retrieval, when the model version changes, and after any incident, with a periodic full exercise as a backstop.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.