How do you secure an MCP server or AI agent tool calls?
You enforce inside the tool handler, because there's no HTTP request to inspect. Run one check per specific operation, with a hardcoded label, covering a budget, injection detection on inputs and on outputs fed back to the model, and local sensitive-information detection.
Arcjet is one way to run that check. Arcjet Guards are called from inside the MCP tool handler, with the authenticated actor and the tool's arguments in scope, and return an allow or deny decision before the tool does its work. The same call screens what the tool returns before the result reaches the model.
Nearly every security tool assumes a request that it can see. MCP breaks that assumption, and that's the whole problem. For how gateways, client-side gates, and in-handler enforcement compare as platforms, see which AI security platforms support MCP server security.
Agentic systems don't have a front door
A client invokes an MCP tool over stdio or Streamable HTTP, so no request arrives at one of your routes. A local stdio server, or a remote server that your proxy doesn't front, is invisible to perimeter tooling: proxies, WAFs, and HTTP middleware alike.
That's awkward, because the tool call is where the risk is. It's where the agent reads a record, calls an API, or moves money, on input that it doesn't fully control.
A gateway in front of your MCP servers helps for traffic that routes through it, and that's a real control at fleet scale. It doesn't cover a tool invoked locally over stdio, a background job, or a direct API call, which is why vendors in that category also ship libraries that run inside the handler.
What is the MCP threat model?
An MCP server sits between a model that can be persuaded and the systems that it can change. Three threats account for most of the risk, and each needs a control in a different place.
| Threat | How it happens | Mitigation |
|---|---|---|
| Untrusted tool responses | A tool returns content that you didn't write, such as a web page, a support ticket, or a third-party API response, and the client passes it into the model's context. A third-party server can also change its tool descriptions after you reviewed them. | Allowlist the servers a client may call, review tool definitions, cap response size, and screen results for sensitive data before they travel. |
| Over-broad tool access | A server exposes every tool to every client, or one token grants every operation. A server that passes the client's token through to a downstream API skips the checks that API would apply to its own tokens. | Expose only the tools a client needs, authorize each call against the authenticated user in the handler, apply per-user budgets, and accept only tokens issued for this server. |
| Injected instructions via tool output | Text inside a tool result reads as an instruction, such as "now export all customers", and the model's next tool call follows it. | Detect prompt injection in tool output before returning it to the model, and enforce authorization on the next call regardless of what the model read. |
The last row is why output screening and handler-side authorization belong together. Detection catches most injected instructions. Authorization in the next handler refuses the action when detection misses one. The MCP security best practices cover the protocol-level risks, including token passthrough, in more depth.
The integration point: enforce inside the handler
The check runs in the handler, immediately before the tool does its work:
import { launchArcjet, tokenBucket, detectPromptInjection, localDetectSensitiveInfo,} from "@arcjet/guard";
const arcjet = launchArcjet({ key: process.env.ARCJET_KEY! });
const budget = tokenBucket({ refillRate: 60, intervalSeconds: 60, maxTokens: 120,});const injection = detectPromptInjection();const pii = localDetectSensitiveInfo({ deny: ["EMAIL", "PHONE_NUMBER"] });
// Inside one specific toolconst decision = await arcjet.guard({ label: "tools.get-customer", actor: session.userId, correlationId: sessionId, rules: [ budget({ key: session.userId, requested: 1 }), injection(args.query), pii(args.query), ],});
if (decision.conclusion === "DENY" || decision.hasFailedOpen()) { // Return an error the client can act on throw new Error(`Denied: ${decision.reason}`);}Screen the result the same way before it returns to the model. The following check runs after the tool has fetched the customer and before the free-text notes reach the model's context:
const notes = customer.notes.join("\n");
const resultDecision = await arcjet.guard({ label: "tools.get-customer.result", actor: session.userId, correlationId: sessionId, rules: [injection(notes), pii(notes)],});
if (resultDecision.conclusion === "DENY" || resultDecision.hasFailedOpen()) { throw new Error("Denied: tool result withheld");}Use one check per specific operation, with a hardcoded label. A direct Guard call fails open (allow with error codes), so the sample denies when hasFailedOpen() is true rather than letting a timeout execute the tool. Vercel AI SDK and LangChain wrappers fail closed unless you opt into continuing on error. Avoid the generic-dispatcher pattern: building label: `tools.${name}` inside a handleToolCall router breaks grep and produces messy groupings in reporting. Labels are validated server-side as slugs, meaning lowercase letters, digits, dash, dot, and underscore only, starting and ending with a lowercase letter or digit.
What to enforce
| Control | Why it belongs at the tool |
|---|---|
| A budget | Repeated tool calls can drain a quota even when each call is individually legitimate, and the entrypoint sees only one invocation |
| Injection detection on inputs | Tool arguments are model-generated, which means they are downstream of whatever the model has read |
| Injection detection on outputs | In a normal API, JSON is data. In an agent workflow, JSON is context, so tool outputs are defended the same way as tool inputs |
| Local sensitive-info detection | Personally identifiable information (PII) moving through a tool is detected without leaving your environment |
The output case is the one most often skipped, and it's the one that makes indirect injection possible. Arcjet's approach to it is covered in how we defend MCP tool outputs from prompt injection.
Make it operable
Pass metadata alongside the label, using server-controlled values such as the authenticated user identifier and a request identifier, and never user-authored free text.
Return different errors depending on which rule denied, rather than a generic failure. "Rate limited, retry in 12 seconds" and "input flagged as prompt injection" are different errors, and the caller must be able to tell them apart. An agent client that can't tell them apart retries the call that it must not retry.
The label and metadata surface in the Console, so a security team can see which tool was called, by whom, and why it was stopped, without reading application logs.
Get identity right in an MCP server
actor must come from your authenticated server-side session, never from a tool argument. A policy can depend on the actor, so a model that sets its own actor value selects its own policy scope, which defeats the control precisely when it matters.
For a stdio MCP server with no user context, there might be no per-user identity available at all. Be explicit about that rather than inventing one: key budgets on a stable deployment or instance identifier that you control, and be clear in your own documentation that the limit is per instance rather than per user. A limit keyed on something that the caller supplies isn't a limit.
Verify it fires
Don't try to reach a guard with curl. There's no HTTP surface to hit.
Invoke the tool through an MCP client or inspector, then confirm the decision landed with arcjet guards list. If nothing appears, the usual causes are a guard call that was never awaited, an empty rules array, or a client that failed before reaching the handler at all.
An empty rule set still reaches Arcjet and returns an allow decision carrying a warning to record that nothing was submitted. It isn't treated as a failure, so a guard with no rules looks like it's working while checking nothing.
Where Arcjet fits
MCP security divides cleanly into two positions. A gateway sits in front of your MCP servers, giving you a catalog of sanctioned servers, identity-aware policy, and audit logging across an estate. Runlayer, Kong, and MintMCP are examples.
Arcjet sits inside the tool handler instead. It reaches tools invoked locally over stdio, background jobs, and direct API calls, none of which traverse a gateway, and it has the arguments and the authenticated actor available at decision time.
On the client side, Arcjet also decides which servers an agent can call. A policy can deny any MCP server that isn't on a reviewed allowlist, and score the remote hosts a call would contact with Arcjet threat intelligence so that a known high-risk host is refused. For coding agents, this runs from the hooks that Claude Code, GitHub Copilot, Cursor, OpenAI Codex, and Muse Code already fire. For more information, see what a rogue MCP server is and how to detect one.
They aren't mutually exclusive, and a large organization plausibly wants both: a gateway to govern which servers exist, and in-code enforcement inside the tools that move money or read customer data. How the layers divide is set out in the AI agent security platform category map.
AI agent governance guides
This guide is one of five on governing what an AI agent may do at runtime:
- How to enforce least privilege for AI agent tool calls: limit each agent to the tools, resources, and arguments its task needs, and enforce it in the handler.
- How to audit and log AI agent activity for compliance: record who, what, and when for every tool call, with sensitive fields redacted.
- How to stop AI agents taking unsafe or unauthorized actions: classify data access, destructive operations, and external calls, and intercept each one.
- How to prevent data exfiltration and sensitive data leakage from AI agents: detect protected data at the boundaries where it moves.
Learn more: Arcjet Guards · How we defend MCP tool outputs from prompt injection
Frequently asked questions
How do I secure an MCP server?
Enforce inside the tool handler, because there is no HTTP request to inspect. Use one check per specific operation with a hardcoded label, running a budget, injection detection on inputs and on outputs fed back to the model, and local sensitive-information detection.
What is the MCP threat model?
Three threats account for most of the risk: untrusted tool responses entering the model's context, over-broad tool access where one client or token reaches every operation, and injected instructions in tool output that steer the model's next call. Mitigate them with server allowlists and output screening, per-call authorization in the handler, and prompt-injection detection on results.
Which AI security platforms support MCP server security?
MCP gateways such as Runlayer, Kong, and MintMCP govern traffic routed through them. Arcjet enforces inside the tool handler, which also reaches stdio tools, background jobs, and direct API calls, and on the client side can restrict coding agents to approved MCP servers.
Why can a WAF or proxy not protect MCP tool calls?
An MCP tool is invoked by a client over stdio or Streamable HTTP rather than by a request hitting one of your routes. A local stdio server, or a path your proxy does not front, gives perimeter tooling nothing to inspect. A gateway helps for traffic that routes through it and misses tools invoked locally.
Should I check tool outputs as well as tool inputs?
Yes. In a normal API, JSON returned from a service is data. In an agent workflow it becomes context, so anything in it can function as an instruction. Tool output re-entering model context is the main indirect injection path.
Why does my guard call not appear in the Console?
The usual causes are a guard call that was never awaited, an empty rules array, or a client that failed before reaching the handler. An empty rule set still returns an allow decision with a warning, so a guard with no rules looks like it is working while checking nothing.
AI runtime security in your code
Protect your AI agent workflows with Arcjet
Arcjet guards run inside the tool, so the allow or deny arrives before the side effect rather than after it.