Leadership · 5 minute read
The CISO's Guide to AI and Agentic AI
A CISO secures agentic AI by treating every agent as a non-human identity acting on untrusted input: scope its permissions to the minimum, gate consequential actions behind approval, defend against prompt injection by design rather than by filtering, prevent data leakage at the gateway, sandbox execution, and keep a tested kill switch for every agent.
An AI agent is software that reads untrusted content and takes actions with real permissions. That single sentence contains the whole security problem. This guide gives CISOs the threat model, the controls that matter most, the questions to ask, and an approach that lets security enable the program instead of becoming the reason it moves into the shadows.
Why are AI agents a new security problem?
Conventional applications execute code you wrote on inputs you validated. An agent executes behavior shaped by a model, on inputs that include documents, emails, tickets, and web pages you did not write, and it has credentials to act. Three properties combine:
- Instructions and data are mixed. A model cannot reliably distinguish content it should process from instructions it should follow.
- The agent has permissions. It can read records, send messages, call APIs, and change state.
- Behavior is probabilistic. The same input can produce different actions, so testing cannot enumerate every path.
The result is a class of attack where a malicious document convinces the agent to use its own permissions against you. FISTA's indirect prompt injection explainer walks through concrete examples.
What does the agent threat model look like?
| Threat | How it happens | Primary control |
|---|---|---|
| Indirect prompt injection | Untrusted content instructs the agent to act | Least-privilege tools; approval gates; isolating untrusted content |
| Data exfiltration | Agent sends sensitive data to a model, a tool, or an external party | Gateway data classification and redaction; egress controls; logging |
| Privilege escalation | Agent uses broad credentials beyond its job | Per-agent identity; per-tool permissions; access reviews |
| Tool and supply-chain compromise | A connector, plugin, or MCP server is malicious or tampered | Signed, reviewed connectors; allowlists; gateway enforcement |
| Runaway or looping behavior | Agent repeats actions or spends without bound | Stop conditions; rate limits; budgets; kill switch |
| Model and prompt tampering | Unauthorized change to prompts, tools, or model configuration | Change control; versioning; evaluation before release |
FISTA's AI agent security architecture whitepaper sets out the reference controls for each row.
Why is prompt injection a design problem?
Filtering malicious instructions out of content does not work reliably, because instructions can be phrased in unlimited ways and models are built to follow them. The effective defense is to assume injection will succeed and limit what it can achieve:
- Least privilege by tool. An agent that summarizes inbound email should not be able to send email or access the CRM.
- Approval gates on consequential actions. Payments, data exports, access changes, and external communications require a person.
- Isolate untrusted content. Process external documents in a context that has no tools, then pass only structured results to the agent that can act.
- Egress controls. Restrict where agents can send data.
With these in place, a successful injection produces a bad summary, not a wire transfer.
What identity and permission model should agents have?
Agents are non-human identities and should be governed like the most sensitive service accounts:
- One identity per agent, issued and revoked through identity governance.
- Permissions granted per tool and per action, at the minimum needed.
- All access through a governed gateway that authenticates, authorizes, and logs every call.
- Periodic access reviews, with the agent's owner recertifying its permissions.
The agent identity and access control whitepaper and the non-human identities for AI agents guide give implementation detail.
How is data leakage controlled?
At the model gateway. Classify data, define which classes may be sent to which models under which terms, redact sensitive fields before calls, log what was sent, and hold providers to contract terms on retention and training use. For workloads that cannot leave the environment, use private deployments. The LLM data loss prevention guide covers the controls in detail.
What operational controls should the CISO require?
- Sandboxing: agents that execute code or browse run in isolated environments with no access to production credentials. See AI agent sandboxing.
- Kill switch: any agent can be disabled in minutes by an on-call engineer, with the procedure tested. See kill switch design.
- Budgets and rate limits: spend and action limits per agent, enforced at the gateway.
- Evaluation and red teaming: adversarial test cases run before release and repeated on a schedule.
- Logging and retention: every run traced and retained for the incident and audit period.
How can security enable rather than block?
Blocking produces shadow AI, which is worse than governed AI. The alternative is a paved road: approved models through the gateway, an agent platform with identity and permissions built in, published risk tiers that define the controls each tier needs, and a security review measured in days. With the road paved, the CISO gains visibility over everything running and can spend attention on the high-risk tier. Security and compliance requirements vary by industry and jurisdiction; this guide is general guidance, not legal advice.
How can FISTA Solutions help a CISO?
FISTA Solutions builds AI agents with per-agent identity, least-privilege tools, approval gates, sandboxing, and kill switches designed in, and helps security teams through its AI enablement practice to define risk tiers, gateway controls, and review processes that scale. Since 2017, FISTA has delivered 150+ projects for 50+ companies across 12+ countries.
If you need a threat model and control set for the agents your business units are already building, talk to FISTA on WhatsApp, or start with the AI agent security risks overview.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What is the biggest security risk of AI agents?
Indirect prompt injection: an agent reads content from an email, document, web page, or ticket that contains instructions, and follows them using its permissions. Because agents act, the consequence is not a bad answer but an unauthorized action such as exfiltrating data or approving a transaction. The defense is bounding what any instruction can cause.
02How should a CISO control what AI agents can access?
Give each agent its own identity, grant permissions per tool at the minimum needed for its job, route all access through a governed gateway that logs every call, gate consequential writes behind human approval, and review agent permissions on the same cycle as user access reviews. Never let agents use shared service accounts.
03How do you prevent data leakage through AI models?
Enforce it at the model gateway: classify data and define which classes may be sent to which models, redact sensitive fields before calls, log what was sent, and contract with providers on retention and training use. For the most sensitive workloads, use private or self-hosted deployments so data never leaves your environment.
04Should security block AI agents until they are proven safe?
Blocking creates shadow AI, which is harder to secure. A better approach is a paved road: approved models through a gateway, an agent platform with identity and permissions built in, a fast security review for new agents, and clear risk tiers that set the controls each tier needs. That lets the CISO enable the program with visibility.
05What should a CISO ask engineering about an AI agent?
What can this agent read, write, and execute? What untrusted content does it process? What happens if that content contains instructions? Which actions require approval? How is it sandboxed? How do we disable it? What is logged, and for how long? What evaluation and red-team testing was done before release, and when is it repeated?
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.