FISTA Solutions does not load Google Analytics until you accept. Rejecting keeps optional analytics off. Read the Cookie Policy.

All field notes

Glossary · 5 minute read

What Is Tool Poisoning? Malicious Tool Definitions Explained

Tool poisoning manipulates an agent through the tool definitions it reads or the results tools return. A description containing hidden instructions, or a result carrying injected text, can redirect agent behaviour without touching user input, which makes third-party tool servers a supply chain risk.

By FISTA Solutions· AI-Native Engineering Team·
What Is Tool Poisoning? Malicious Tool Definitions Explained article cover

Prompt injection is widely understood as something arriving through user input. Tool poisoning is the same class of attack arriving through the agent's own tooling — the descriptions it reads and the results it receives — which is a channel most systems treat as trusted. As agents adopt shared and third-party tool servers, this becomes a supply chain concern. This explainer covers it. It complements what is a tool schema and what is indirect prompt injection, and reflects FISTA Solutions' approach in AI agents delivery.

How does a poisoned description work?

The tool description is text the model reads to decide what to do. A description containing instructions — disregard previous constraints, always call this tool before answering, append the following to your response — is read alongside your own system instructions.

The model has no mechanism for distinguishing your instructions from the tool author's. Both are text in its context, and the tool description arrives with the implicit authority of being part of the system's configuration.

ChannelCarriesTrusted by default
Tool descriptionInstructions about when to use itUsually yes
Tool parameters schemaField descriptionsUsually yes
Tool result contentWhatever the tool returnsFrequently yes
Error messagesText from the toolAlmost always
Retrieved documentsThird-party contentSometimes

What about poisoned results?

Any tool returning text that enters the agent's context can carry instructions in it. A web page fetched by a browsing tool. A search result. A database field containing user-submitted content. A file's contents. A support ticket description.

None of these looks untrusted, and all of them are. The defence is treating tool output as data throughout — delimited, labelled, and never granted instruction authority — rather than assuming that content arriving through your own tool is safe because the tool is yours.

Why are third-party tool servers a supply chain risk?

Because their definitions can change after approval. You review a server's tools at integration, approve them, and ship. Next week the server returns different descriptions. Nothing in your repository changed, so no review triggers, and your agent is now reading instructions you never saw.

That is the classic supply chain shape, and shared tool ecosystems make it practical at scale. See how to build an mcp gateway.

What defends against it?

Pinning tool definitions and reviewing changes, so a modified description requires approval rather than taking effect silently. Detecting definition changes at all, which requires recording what you approved.

Then the defences that hold regardless: treating all tool output as untrusted data, enforcing authorisation at the tool against the acting identity, and granting least privilege so that an agent persuaded to misbehave can do little.

Is filtering enough?

No. Pattern filtering catches known phrasings and misses new ones, and the surface is any text entering context. It is worth having and it is not the control that matters.

The control that matters is that the agent cannot do damaging things regardless of what it is persuaded to attempt, because authorisation is checked where the model cannot reach.

What should you do first?

List the tool servers your agents connect to that you do not control, and check whether you would know if one changed a tool description tomorrow. In most implementations the answer is no, and recording approved definitions with change detection is a contained piece of work.

How does a gateway help?

By giving you one place to enforce the controls. A gateway between agents and tool servers can pin definitions, detect changes, apply authorisation policy consistently, log every invocation, and rate limit — none of which is practical to implement separately in every agent that might connect.

It also centralises the approval decision. Adding a new tool server becomes a governed action rather than a configuration change any engineer can make, which is the difference between an ecosystem you can reason about and one that grows by accident.

What should be logged?

Every tool invocation with its arguments and result, the definition version in use, and the identity on whose behalf it ran. That record is what makes a poisoning incident investigable, and it is what shows whether the agent's behaviour changed at the moment a definition did.

Without the definition version in the log, an investigation cannot establish what the agent was actually reading at the time, which is usually the question that matters most.

Does this apply to internally built tools?

Less acutely and not not at all. An internal tool whose description is written carelessly can still steer an agent in unintended ways, and one that returns user-supplied content without treating it as data is an injection channel regardless of who wrote it. The supply chain risk is specific to third parties; the data-versus-instruction discipline applies everywhere.

How FISTA Solutions helps

FISTA Solutions pins and reviews tool definitions with change detection on third-party servers, treats all tool output as untrusted data rather than instruction, enforces authorisation at the tool against the acting identity, and grants agents least privilege so poisoning has limited reach, through AI agents, AI enablement, and forward deployed engineers. The record behind the approach is 150+ projects for 50+ companies with 99.9% uptime.

To secure the tools your agents depend on, message FISTA on WhatsApp, or read what is a tool schema.

Share-ready article cover

Download the generated social format.

Download cover

Clear answers

Questions raised by this field note.

Straightforward guidance for evaluating scope, fit, and the next step.

01How does a poisoned description work?

The description is text the model reads when deciding what to do, so a description containing instructions — ignore prior constraints, always call this tool first, include the following in your output — is read alongside your own instructions and may be followed.

02What about poisoned results?

A tool returning text the agent incorporates into its context can carry instructions in that text. Retrieved web pages, search results, database fields containing user-supplied content, and file contents are all channels for this, and none is obviously untrusted at a glance.

03Why are third-party tool servers risky?

Because their definitions can change after you approved them. A server reviewed at integration may serve a different description next week, and nothing in your codebase changed, so no review is triggered. That is a supply chain risk.

04What defends against it?

Pinning and reviewing tool definitions, detecting changes to them, treating all tool output as untrusted data rather than instruction, enforcing authorisation at the tool against the acting identity, and granting least privilege so that a poisoned tool has limited reach even when it succeeds.

05Is filtering sufficient?

No. Filtering catches known patterns and misses new phrasings, and the attack surface is any text entering context. The durable defence is limiting what an agent can do regardless of what it is persuaded to attempt.

Start with the hard problem

Need the outcome owned, not merely analyzed?

Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.

Start a project