How to stop AI agents from taking unsafe or unauthorized actions

Classify every agent tool as data access, destructive, or external, then refuse each call by default at runtime.

8 min read
In short: Classify each tool an agent can call as data access, a destructive operation, or an external call, and intercept every call at runtime in the code that executes it, refusing by default. Arcjet Guards send typed facts about the call to a policy that allows or denies before the side effect.

How do I stop AI agents from taking unsafe or unauthorized actions?

Classify every tool an agent can call by what it does: reading data, destroying or changing something that is hard to undo, or sending something outside your system. Then intercept each call at runtime, in the code that executes the tool, and refuse it unless the authenticated user, the target resource, and the arguments pass the rules for that class. Refusing by default is what stops an unsafe action. Asking the model to behave doesn't.

Arcjet is one way to run that interception. An Arcjet Guard call in the tool path sends typed facts about the call, such as its class, the user's role, and the destination host, to a policy that returns allow or deny before the side effect runs. A security team can change the policy without an application deploy, and every decision is recorded with its reason.

The rest of this guide covers the three classes, where to intercept, and a worked policy that you can adapt.

Classify unsafe actions into three kinds

Most unsafe agent actions fall into one of three classes. Each has a different failure mode, so each needs a different check.

ClassExamplesWhat goes wrongCheck before it runs
Data access

Read a customer record, query a warehouse, search a document store

Another tenant's data, more records than the task needs, or fields the user can't see

Tenant-scoped lookup, record-count limit, and field-level permission

Destructive operationsDelete a project, issue a refund, change production configurationAn irreversible change that the user didn't intend or can't makeRole check, bounded amounts, and a recorded approval
External calls

Send an email, call a webhook, fetch a URL, post to a chat channel

Data leaving your system, or a request to an attacker's hostDestination allowlist, recipient binding, and content screening

Keep the classification on the server, in a table the model can't change. A tool that isn't in the table is denied. An "unknown" class is a bug to fix, not a default to allow.

Some tools span classes. An export that reads a thousand records and emails them is both data access and an external call, so apply both sets of checks. The lethal trifecta describes why that combination is the one to design against.

How do I stop AI agents from accessing data they shouldn't?

Resolve every resource from the authenticated session, not from the model's arguments. If the handler looks up invoiceId scoped to the session's tenant, then another tenant's invoice doesn't exist from the agent's point of view, whatever ID the model supplies.

Then bound the reach. Cap the number of records a single call can return, return only the fields the user may see, and deny bulk reads the task doesn't need. Scoped credentials limit what the agent could reach in the worst case. A runtime check limits what it does with that reach on this call. For both halves in detail, see stopping AI agents accessing data they shouldn't.

Hold destructive operations to a stricter bar

A destructive operation needs a stronger reason to allow than a read. Check the user's role for this operation, bound amounts with server-side values, and require an approval that is bound to the exact resource and arguments. If the model changes the amount after approval, evaluate the new action again.

Make the operation idempotent with an operation ID, so a retry doesn't delete or refund twice. For the full pattern, including why a human click is a hold rather than a policy, see preventing irreversible AI agent actions.

Restrict where external calls can go

An external call is how data leaves and how an agent reaches a host an attacker controls. Allowlist destination hosts in a policy the agent can't edit, bind recipients to the verified case or account, and screen outbound content for sensitive data. See egress allowlists for AI agents and detecting malicious URLs.

Intercept at runtime, at the point the tool executes

Put the check in the code that performs the action, immediately before the side effect. That's the only point every path shares: the model's direct call, a retry, a background worker, and an alternate tool that reaches the same service.

Checks elsewhere leave gaps. A system prompt can be overridden by injected text. A tool-name allowlist in the client doesn't check arguments. A gateway sees only the calls routed through it. Keep those layers, but don't count them as the enforcement point. For a comparison of the layers, see least privilege for AI agent tool calls.

When the check can't complete, refuse the destructive and external classes. A timeout in the policy service mustn't become permission to delete.

Example: a policy for each class of action

Give each tool its own guard call with a hard-coded label, and send typed facts about the call that the policy for that label can read. The following destructive tool resolves the project within the user's tenant, then asks Arcjet whether this user may delete it:

import { launchArcjet, policyInput } from "@arcjet/guard";
const arcjet = launchArcjet({ key: process.env.ARCJET_KEY! });
export async function deleteProject(
args: { projectId: string },
session: { userId: string; tenantId: string; role: string },
) {
// Tenant-scoped lookup: another tenant's project doesn't exist here.
const project = await projects.find({
id: args.projectId,
tenantId: session.tenantId,
});
if (!project) return { error: "Not permitted" };
const approval = await approvals.find({
resourceId: project.id,
operation: "delete",
});
const decision = await arcjet.guard({
label: "tools.delete-project",
actor: session.userId,
inputs: {
role: policyInput.server.string(session.role),
environment: policyInput.server.string(project.environment),
approved: policyInput.server.boolean(approval?.status === "approved"),
},
});
if (decision.conclusion === "DENY" || decision.hasFailedOpen()) {
return { error: "Delete denied by policy" };
}
return projects.delete(project.id);
}

A Guard policy for the tools.delete-project label then states the rules. It reads only the values the application sent:

package arcjet.guard
import rego.v1
deny contains "requires-admin" if {
input.values.role != "admin"
}
deny contains "production-requires-approval" if {
input.values.environment == "production"
not input.values.approved
}

The other two classes follow the same shape, each on its own label. A tools.get-customer policy bounds a data-access tool by the number of records it would return, and a tools.send-webhook policy restricts an external call to reviewed hosts:

deny contains "bulk-read" if {
input.values.record_count > 100
}
deny contains "unlisted-destination" if {
not input.values.destination_host in {"api.stripe.com", "hooks.slack.com"}
}

Three properties make this safe to run. Every input comes from server state, so the model can't claim to be an admin, approve its own request, or supply its own allowlist. The tenant check stays in code, where the data is. A policy starts in dry run, so you can review what it would have denied against real traffic before it blocks, and a live rule needs a stored test before it can publish. For the authoring workflow, see Guard remote policies.

Add rules to the same call for checks that aren't expressions over inputs, such as a per-user rate limit on destructive actions or prompt-injection detection on a free-text argument.

Test that unsafe actions are refused

For each class, start from an allowed call, then change one fact at a time: a non-admin role, a production project with no approval, 101 records, an unlisted host, and another tenant's resource ID. Assert that the decision is a deny and that the downstream call count is zero. Run the same cases through retries and background workers that reach the same services.

Include the adversarial case. Put an instruction in a tool result or a retrieved document that asks for a destructive action, and confirm that the policy refuses it even though the model proposed it. See AI agent security testing for building the full matrix.

AI agent governance guides

This guide is one of five on governing what an AI agent may do at runtime:

Learn more: Arcjet Guards · Guard remote policies

Frequently asked questions

How do I stop AI agents from taking unsafe or unauthorized actions?

Classify each tool as data access, a destructive operation, or an external call in a server-owned table. Intercept every call in the code that executes it, resolve the target from the authenticated session, and refuse by default unless the user, resource, and arguments pass the rules for that class.

How do I stop AI agents from accessing data they shouldn't?

Resolve every resource from the session's tenant rather than from the model's arguments, cap the records and fields a call can return, and scope the credential. Check what the tool returns before it travels into model context or a response.

Where should I intercept an AI agent's actions?

In the code that performs the action, immediately before the side effect. That point is shared by direct calls, retries, background workers, and alternate tools; a system prompt, client allowlist, or gateway each misses some of them.

How does Arcjet stop unsafe agent actions?

A guard() call in each tool sends typed, server-derived inputs, such as the user's role, an approval status, or a destination host, to the Guard policy for that tool's label, written in Rego. The policy returns allow or deny before the side effect, a security team can change it without a deploy, and each decision is recorded with its reason.

AI runtime security in your code

Protect your AI agent workflows with Arcjet

Keep the enforcement point in code you review, and let an authorized operator change the policy centrally without an application deploy.