What is session tainting for AI agents?

Mark an agent session untrusted after a risky finding, and lock risky follow-up actions until a person reviews it.

7 min read
In short: Session tainting marks an AI agent's session as untrusted after something risky happens in it, such as a prompt injection in a tool result or credentials touched, and then refuses shell, web, MCP, and file-write calls for the rest of the session until an administrator clears it with a reason. It covers the gap per-call policies leave: content the agent has already read can't be withheld, but what it does next can be stopped.

What is session tainting for AI agents?

Session tainting marks an AI agent's session as untrusted after something risky happens in it, and then restricts what the agent can do for the rest of that session. A finding triggers the taint, such as a prompt injection in a tool result, or sensitive data exposed in a file the agent read. Once a session is tainted, later tool calls – shell commands, web requests, MCP calls, and file writes – are refused until an administrator reviews the session and clears the taint.

The idea borrows from taint tracking in programming languages, where data from an untrusted source is marked so that it can't reach a sensitive operation without being checked. Session tainting applies the same rule to a whole agent session: once untrusted content has entered the agent's context, the session can't take risky actions.

Why do AI agents need session tainting?

Most agent controls judge one action at a time. A policy on a tool call asks whether this command, file read, or fetch is allowed, based on what's in the call. That's the right check for most rules, and it has one blind spot: it can't see what the agent has already read.

Consider a coding agent that fetches a web page as part of a task. The page contains hidden instructions to collect credentials and send them to an external host. By the time anything can inspect the page, the fetch has already run and the text is in the model's context. There's no way to take it back. The next step might be an ordinary-looking curl that no per-call rule would object to.

Session tainting closes that gap. The finding in the fetched page doesn't block the fetch – it can't – but it changes what the rest of the session is allowed to do. The agent can keep reading files and answering the developer, and it can't run commands, write files, or make network requests until a person has looked at what happened.

How does session tainting work?

Session tainting has three parts:

  1. A finding sets a flag. Something screens the session's activity, such as the text a tool returned, and sets a named flag on the session when it finds a risk.
  2. Policies read the flags. Each later tool call carries the session's active flags, and a policy denies risky kinds of call while a flag is set.
  3. A person clears the flag. The session stays locked until an administrator reviews it and removes the taint. The agent can't clear it itself.

In Arcjet, the flags are set by Arcjet, never by the agent, and they arrive on each tool call as the session_flags policy input. The following flags can taint a session:

FlagSet when
prompt_injectionA tool result looks like an instruction override or jailbreak
sensitive_data_exposure

A tool result contains an entity type the policy denies, such as a card number or Social Security number

credential_access

Session activity touched secrets, keys, tokens, or credential paths

data_exfiltration

Local files or secrets were sent to an external host. On shell tools, this also requires upload evidence in the command, so an ordinary package install doesn't lock the session

code_execution_risk

Untrusted code was fetched and run, piped into a shell, or used to escalate privileges

elevated_riskThe session as a whole crossed a high risk threshold

The first two flags come from screening tool results with Arcjet's prompt injection and sensitive information detectors. The rest come from Arcjet's risk scoring of the session's sequence of actions. Claude Code, GitHub Copilot, Cursor, and OpenAI Codex all report tool results through a post-tool-use hook, which is where the screening runs.

What does a session taint policy look like?

The coding-agent.tainted-session starter policy in Arcjet locks shell, web, MCP, and file-write calls while any of the flags is active, and still allows reads:

risky := {"shell", "web", "mcp", "file_write"}
lock_flags := {
"prompt_injection",
"credential_access",
"data_exfiltration",
"code_execution_risk",
"sensitive_data_exposure",
"elevated_risk",
}
deny contains "tainted-session-risky" if {
some flag in input.values.session_flags
flag in lock_flags
input.values.tool_kind in risky
}

You tune the policy by editing the two sets. For example, to lock only network access after an exfiltration finding, narrow risky to {"web", "mcp"} and lock_flags to {"data_exfiltration"}. The same policy declares which tool-result screens run: its prompt injection detector enables the prompt_injection flag, and its sensitive information detector, with the entity types to deny, enables sensitive_data_exposure.

How is a tainted session cleared?

In Arcjet, only a person in the Arcjet Console can clear a taint, and clearing it requires a written reason. The clear is recorded in the audit log, so every unlocked session has a named reviewer and an explanation. The agent can't clear its own taint through the hook or through MCP, and a SessionEnd event doesn't clear it either.

That review step is the point of the design. The finding that set the flag might be a false positive, in which case the reviewer clears it and the developer carries on. Or it might be the start of an incident, and the session's recorded activity shows exactly what the agent read and tried to do next.

What doesn't session tainting do?

  • It doesn't undo the action that triggered it. A tool result has already reached the model when it's screened. Tainting stops what comes next; it can't withhold what the agent has read.
  • It isn't set by a denied tool call. A policy that blocks the AWS CLI doesn't taint the session on its own. Flags come from tool-result screening and session risk scoring. To stop a specific action, use a per-call policy. For an example, see stopping coding agents using AWS credentials.
  • It doesn't count events. A rule such as "taint after 10 commands in auto mode" isn't something a taint policy expresses, because flags come from findings, not from counters. To control auto mode, use a policy on the permission mode. For more information, see requiring human approval in Claude Code auto mode.
  • It only runs the screens you publish. Tool-result screening runs only when a published policy reads session_flags and declares the matching detector.

How does session tainting compare with other controls?

ControlDecides based onGood at
Per-call policyThe current tool call's command, paths, and destinationsStopping a known-bad action, such as reading SSH keys
Human approvalA person reviewing each risky actionHigh-stakes actions where a person can judge intent
Session taintingFindings earlier in the same session

Stopping follow-up actions after untrusted content enters the context

Kill switchAn operator decision to stop an agent entirelyIncident response once a problem is confirmed

The controls work together. Per-call policies stop the actions you can name in advance, session tainting contains the ones that follow an attack you couldn't prevent, and review turns each locked session into a decision a person made. For how these layers fit a coding agent rollout, see coding agent security tools compared.

Frequently asked questions

What is session tainting?

Session tainting marks an AI agent's session as untrusted after a risky finding, such as a prompt injection in a tool result, and then restricts later actions in that session – typically shell commands, web requests, MCP calls, and file writes – until a person reviews the session and clears the taint.

Why do AI agents need session tainting?

Per-call policies judge each action on its own inputs, so they can't see what the agent has already read. Once an injected web page or tool result is in the model's context it can't be withdrawn, and the next command may look ordinary. Tainting changes what the rest of the session is allowed to do.

What triggers a session taint in Arcjet?

Arcjet sets flags such as prompt_injection and sensitive_data_exposure by screening tool results, and credential_access, data_exfiltration, code_execution_risk, and elevated_risk from risk scoring of the session's actions. The agent never sets flags itself.

How is a tainted session cleared?

In Arcjet, only a person in the Arcjet Console can clear a taint, with a required reason that is recorded in the audit log. The agent can't clear it through the hook or MCP, and a SessionEnd event doesn't clear it.

What doesn't session tainting do?

It doesn't undo the tool result that triggered it, it isn't set by a denied tool call on its own, it doesn't count events, and it only runs the tool-result screens that a published policy declares.

AI runtime security in your code

Protect your AI agent workflows with Arcjet

Publish the tainted-session starter policy in dry run and see which follow-up calls it would lock.