What is MCP tool poisoning?
MCP tool poisoning is an attack in which an adversary manipulates the descriptions, schemas, metadata, definitions, or responses of a Model Context Protocol (MCP) tool so an AI agent behaves differently than the user intended.
OWASP describes tool poisoning as an attack where the contract between an agent and a tool becomes untrustworthy. The attacker may change how a benign-looking action maps to real behavior, hide instructions inside tool metadata, or return content designed to manipulate the model after the tool is called.
The important point is that the agent can remain protocol-compliant while still taking an unintended action.
How MCP tool poisoning works
A typical agent sees a catalog of tools and uses their names, descriptions, and parameter schemas to decide what to call and how. An attacker can abuse that trust in several ways.
- Poisoned descriptions: Embed hidden instructions that influence l the model to access additional data or call another tool.
- Schema poisoning: Changes parameter definitions so a harmless-looking operation maps to a destructive action or a broader scope than expected.
- Poisoned responses: Return legitimate-looking data mixed with instructions that the model interprets as trusted context.
- Definition changes after approval: Let a server present one behavior during review and change the tool later, creating a form of rug-pull attack.
The attack is especially risky when untrusted MCP tools share an agent context with privileged tools such as filesystem access, internal APIs, or databases.
Why traditional API security is not enough
API security assumes the caller knows what endpoint it is invoking and why: it focuses on authenticating the caller and protecting the request and response. With an AI agent, the model may dynamically choose the tool to call and how to use it based on natural-language metadata.
That means the metadata itself becomes part of the security boundary. Authentication and TLS can confirm that the agent is connected to the expected server, but they do not prove that the server’s tool description is trustworthy or that a tool response is safe to place directly into the model context.
How to reduce MCP tool poisoning risk
- Approve MCP servers and tools deliberately: Do not allow agents to connect to arbitrary servers by default.
- Treat descriptions and responses as untrusted input: Do not assume tool-provided text is safe simply because it arrived through an approved MCP connection.
- Constrain privileged tools: Restrict privileged tools so high-impact capabilities have server-side authorization and policy checks an injected instruction cannot override.
- Use structured responses: Where practical, use structured responses since fixed schemas reduce the amount of free-form text that can carry hidden instructions.
- Fingerprint tool definitions: Detect material changes in schemas, descriptions, and capabilities after a tool was approved.
- Require human approval: Enforce human approval for high-impact actions such as deletion, code deployment, financial actions, or broad data export, with controls outside the model’s own instruction context.
Tool poisoning vs. prompt injection
MCP tool poisoning is closely related to indirect prompt injection, but the attack surface is different and more specific.
- A traditional indirect prompt injection hides malicious instructions in content the model reads, such as a document or web page.
- MCP tool poisoning places the manipulation inside the tool relationship itself, including its metadata, schema, or returned content. That gives the attacker a trusted-looking channel directly into the agent’s decision process.
Both can influence an agent’s decision-making, but MCP tool poisoning specifically targets the trust placed in MCP tools and their associated information.
FAQs
1. Is MCP tool poisoning the same as a malicious MCP server?
A malicious MCP server can use tool poisoning, but tool poisoning describes the manipulation technique. A server may also be compromised after originally being legitimate and subsequently delivering poisoned metadata or responses.
2. Can tool poisoning happen after a tool was approved?
Yes. A server can change its tool definitions or responses after review, which is why ongoing fingerprinting and change monitoring matter.
3. Does authentication prevent MCP tool poisoning?
No. Authentication can verify who the server is, but it does not guarantee the server’s tool metadata, schema, or output is safe.
4. What is the most important control for high-risk MCP tools?
Enforce critical restrictions outside the model. Backend authorization, scoped permissions, and explicit approval gates should prevent an injected instruction from gaining additional privileges.