FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Playbook ┬╖ 6 minute read

How to Run an AI Red Team Exercise That Finds Real Issues

An AI red team exercise finds real issues when it is scoped to a specific system with a defined objective, staffed with domain people as well as security ones, covers both prompt-level and authority-level attacks, and converts findings into controls and test cases rather than a report.

By FISTA Solutions┬╖ AI-Native Engineering Team┬╖
How to Run an AI Red Team Exercise That Finds Real Issues article cover

AI red teaming finds different problems from conventional security testing, because the interesting failures are behavioural rather than infrastructural. This playbook covers running an exercise that produces controls rather than a document, drawing on FISTA Solutions' AI agents work.

When is this worth doing?

Before a system with meaningful authority or sensitive data reaches production, and after significant changes to its tools, model, or scope.

It is not worth doing on a prototype nobody has decided to ship. The exercise costs real time from scarce people, and running it on something that may be abandoned wastes both.

What does the sequence look like?

StepPurpose
1. Scope and objectiveOne system, a stated goal
2. Assemble the teamSecurity, domain, architecture
3. Establish rulesWhat is in bounds, what is recorded
4. Run the categoriesInjection, authority, exfiltration, behaviour
5. Record findings properlyReproducible, severity, evidence
6. Convert to controlsTests, restrictions, accepted risks

Step 1 тАФ Scope to one system with an objective

Pick a single system and state what the exercise is trying to establish: whether the agent can be made to take an unauthorised action, whether the assistant can be made to reveal restricted information, whether the classifier can be manipulated.

Broad exercises across an estate produce shallow coverage. A focused exercise against one system with a clear objective produces findings specific enough to act on.

Write the objective down. Exercises without one drift towards whatever the participants find interesting, which is rarely what matters most.

Step 2 тАФ Assemble a team with domain knowledge

Security people know how to attack systems. Domain experts know what a harmful output looks like in your context тАФ a wrong clinical statement, an unauthorised price commitment, a discriminatory phrasing тАФ and security specialists frequently cannot recognise those.

Include someone who understands the architecture too. Knowing which tools the agent has and what they can reach directs the effort towards the attacks that matter rather than the ones that are fun.

Three to five people for a day or two produces more than a large group for an afternoon.

Step 3 тАФ Establish the rules before starting

Agree what is in bounds, what environment is being used, what data is real, how findings are recorded, and who is told what.

Use a non-production environment with realistic data where possible. Red teaming a live system with real customers risks causing the harm you are testing for, and using synthetic data that does not resemble reality produces findings that do not transfer.

Agree the disclosure path too. Findings that are genuinely serious need a route to someone who can act on them the same day.

Step 4 тАФ Run the attack categories deliberately

Cover prompt injection through untrusted content the system reads, authority abuse where the agent is induced to take actions outside its scope, data exfiltration through outputs, and behavioural failures in high-consequence cases.

Authority abuse is the category that matters most and gets the least attention. An agent that can be talked into calling a tool it should not is a more serious problem than one that can be made to write something impolite. See what is a confused deputy attack.

Test through the paths real content arrives on: documents, emails, tickets, web pages. Injection through a pasted prompt is the least realistic delivery mechanism and the most commonly tested.

Step 5 тАФ Record findings so they are reproducible

Each finding needs the exact input, the observed output or action, the conditions, a severity assessment, and a note on what made it possible.

Reproducibility is what separates a finding from an anecdote. A report saying the agent 'sometimes' took an unauthorised action produces argument; one with the exact sequence produces a fix.

Severity should reflect consequence rather than cleverness. An easily-triggered minor issue usually matters more than an elaborate attack requiring conditions nobody will meet.

Step 6 тАФ Convert findings into controls

Every finding becomes one of three things: a control that prevents it, a regression test that catches it recurring, or an accepted risk with a named owner and a reason.

That conversion is the whole point. Exercises that end with a report get run once, because the second one finds the same issues and everyone concludes the exercise is not useful.

Prefer structural controls over instructional ones. Restricting what a tool can do is durable; adding an instruction telling the model not to do something is not, and the difference becomes obvious at the next exercise. See what is least privilege for ai agents.

What about the model provider's safeguards?

They help and they are not your control. Provider safety measures change without notice, vary by model version, and are designed for general harms rather than your specific risks.

A system whose safety depends entirely on the provider has outsourced a control it cannot verify or version. Test what happens when those safeguards do not fire, because eventually a model update will change where they do.

How do you avoid theatre?

By measuring whether findings led to changes. An exercise that produced twelve findings and zero code changes was theatre regardless of how thorough it looked.

Track the conversion rate from finding to control, and report it. Organisations where that number is low do not have a red teaming problem; they have a prioritisation problem, and naming it honestly is more useful than running another exercise.

Who needs to be involved?

Two to four people for a focused exercise: security, domain, and architecture, with a facilitator keeping it to the objective.

External participants bring useful unfamiliarity. People who did not build the system ask questions its builders stopped asking.

How long does it take?

One to three days of concentrated effort for a focused exercise, plus a week for writing up and converting findings. Longer exercises produce diminishing returns without broader scope.

What are the common failure modes?

Scoping across an estate. No domain experts. Testing only prompt-level attacks. Findings recorded as anecdotes. Fixes that are instructions rather than restrictions. And no repeat after changes.

How do you know it worked?

Findings converted into controls and tests, the same issues not recurring at the next exercise, and the team able to state what the system will refuse and why.

What does it cost?

Mostly people's time rather than tooling. The expensive version is the one that stalls halfway and leaves the organisation with neither the old state nor the new one, which is why a narrow first pass beats a comprehensive plan nobody finishes.

Budget the work as an operated change rather than a project with an end date, because most of these need a maintenance tail. See AI total cost of ownership.

What should you do first?

Pick one system and list every tool it can call. Ask whether any content it reads could persuade it to use one inappropriately. That question alone usually produces the first finding.

How FISTA Solutions helps

FISTA Solutions runs this work alongside client teams rather than around them: exercises scoped to a system with a stated objective, findings converted into structural controls and regression tests rather than a report, evidence produced as the work proceeds, and handover that leaves your people able to continue without us. Delivery runs through AI agents, AI enablement, and forward deployed engineers. The record is 150+ projects for 50+ companies across 12+ countries, with 47% average efficiency gains where measured.

To run this with support, message FISTA on WhatsApp, or read AI red teaming cost.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01How is AI red teaming different from security testing?

It targets behaviour rather than infrastructure. The interesting failures are a system persuaded to do something it should not, act outside its authority, or reveal information it holds, rather than an unpatched dependency.

02Who should be on the team?

Security people, domain experts who know what a harmful output looks like in your context, and someone who understands the system's architecture. Domain experts find issues security specialists cannot recognise.

03What attack categories matter most?

Prompt injection through untrusted content, authority abuse where the agent takes actions it should not, data exfiltration through outputs, and behavioural failures such as confidently wrong answers in high-consequence cases.

04What happens to the findings?

Each becomes a control, a regression test, or an accepted risk with a named owner. Findings that become a report and nothing else are the usual outcome and the reason exercises get run once.

05How often should this run?

After significant changes to the system, its tools, or its model, and at least periodically. A single exercise describes the system as it was that week, and agents change more often than that.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. WeтАЩll map the fastest credible path from intent to verified production.

Start a project