Checklist · 4 minute read
AI Penetration Test Checklist: Scoping a Test That Finds Things
An AI penetration test looks for different things from an application test: instructions smuggled through content, tool capabilities chained into unintended paths, data extracted through the model, and authorisation bypassed by persuasion. Scope those explicitly or the test will report conventional findings only.
AI penetration tests find different things from application tests, and only if the scope says so. This checklist covers scoping one, drawn from FISTA Solutions' AI agents security work.
What attack classes should be in scope?
Six classes specific to these systems.
| Class | What it tests |
|---|---|
| Direct injection | Instructions in user input |
| Indirect injection | Instructions in retrieved content |
| Capability chaining | Tools combined into a path |
| Authorisation bypass | Persuading past a control |
| Data extraction | Pulling data through the model |
| Denial and cost abuse | Driving spend or exhausting limits |
Scope definition
Say what is in scope or you will get a conventional report.
- AI-specific attack classes named explicitly in the scope
- Systems and agents in scope listed
- Tools and integrations in scope listed
- Retrieval sources the tester may place content in identified
- Production versus staging decided and stated
- Test accounts and data provided
- Success criteria for the engagement agreed
Injection testing
Both direct and indirect. See why agent security is different.
- Direct injection through user input tested
- Indirect injection through documents tested
- Injection through tickets, email, or other ingested content tested
- Injection through web content the system fetches tested
- Multi-turn and delayed injection tested
- Encoded and obfuscated payloads tested
- Non-English injection tested where relevant
Capability and authorisation
Where agents are concerned, this is the highest-consequence area.
- Each tool tested for parameter abuse
- Capability combinations tested for unintended paths
- Authorisation checks tested by persuasion attempts
- Limit enforcement tested by attempting to exceed
- Delegation boundaries tested if the agent delegates
- Actions attempted outside the documented scope
- Credential scope tested for over-permissioning
Data exposure
What can be pulled out through the model.
- Attempts to extract other users' data
- Attempts to extract system prompts and configuration
- Attempts to retrieve documents the user cannot access
- Cross-tenant isolation tested where multi-tenant
- Logs and error messages checked for leakage
- Cached responses checked for cross-user exposure
- Training or example data leakage tested
Availability and cost
Frequently omitted and genuinely exploitable.
- Requests designed to maximise token consumption tested
- Agent loops driven toward the step limit
- Rate limit behaviour under abuse tested
- Cost controls tested by attempting to exceed them
- Resource exhaustion through large inputs tested
- Downstream system exhaustion through tool calls tested
- Recovery after abuse observed
Engagement management
Agents take real actions; the rules need to be explicit.
- Actions that must not be executed listed
- Stop procedure agreed and reachable
- Point of contact available throughout
- Monitoring in place to observe the test
- Data touched during the test handled and deleted
- Findings reported with reproduction steps
- Retest scheduled after remediation
What are the most common failures?
Scoping it as a web application test. Omitting indirect injection. Testing tools individually. Ignoring cost abuse. And engaging testers without AI-specific experience.
Who should own this?
Security owns the engagement; the system owner agrees the boundaries and accepts residual risk after remediation. Findings without an owner do not get fixed.
How often should it run?
Before launch for systems with meaningful consequence, after any significant capability change, and annually thereafter. Retest after remediation rather than accepting a fix report.
What evidence should it produce?
The scope document, the findings report with reproduction steps, remediation records, and retest results. That set demonstrates the loop closed.
What if you cannot afford external testing?
Run internal adversarial testing against the same list. It is weaker than independent testing and considerably better than nothing.
Build the attack cases into the evaluation suite so they run on every change. That converts a one-off exercise into a continuous control, which is the more valuable outcome anyway. See AI agent security risks.
What should you do first?
Place a document containing instructions where your system will retrieve it and see what happens. That single test is the most informative one available.
How FISTA Solutions helps
FISTA Solutions builds and operates production AI systems through AI agents, AI enablement, and forward deployed engineering: AI-specific attack classes named explicitly in scope, with indirect injection and capability chaining tested rather than conventional application findings, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To adapt this checklist to your environment, message FISTA on WhatsApp, or read agent permission review checklist.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01How does this differ from an application test?
The attack surface includes content the system reads. A tester places instructions in a document or a ticket rather than crafting a request, which conventional application testing does not cover.
02What is indirect injection?
Instructions placed in content the system will retrieve â a document, a web page, an email â rather than sent directly. It is the highest-value class to test because the attacker needs no access to your interface.
03What is capability chaining?
Combining individually permitted tool calls into an unintended outcome, such as reading sensitive data and then sending it externally. Tool-by-tool review does not find these.
04What should testers have?
Experience with model-specific attack classes, not just web application testing. A conventional tester will produce a conventional report that misses the AI-specific exposure.
05What boundaries are needed?
What data may be touched, whether production is in scope, what actions must not be executed, and how to stop. Agents can take real actions, so the rules need to be explicit.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.