What is session tainting for AI agents?
Session tainting marks an AI agent's session as untrusted after something risky happens in it, and then restricts what the agent can do for the rest of that session. A finding triggers the taint, such as a prompt injection in a tool result, or sensitive data exposed in a file the agent read. Once a session is tainted, later tool calls – shell commands, web requests, MCP calls, and file writes – are refused until an administrator reviews the session and clears the taint.
The idea borrows from taint tracking in programming languages, where data from an untrusted source is marked so that it can't reach a sensitive operation without being checked. Session tainting applies the same rule to a whole agent session: once untrusted content has entered the agent's context, the session can't take risky actions.
Why do AI agents need session tainting?
Most agent controls judge one action at a time. A policy on a tool call asks whether this command, file read, or fetch is allowed, based on what's in the call. That's the right check for most rules, and it has one blind spot: it can't see what the agent has already read.
Consider a coding agent that fetches a web page as part of a task. The page contains hidden instructions to collect credentials and send them to an external host. By the time anything can inspect the page, the fetch has already run and the text is in the model's context. There's no way to take it back. The next step might be an ordinary-looking curl that no per-call rule would object to.
Session tainting closes that gap. The finding in the fetched page doesn't block the fetch – it can't – but it changes what the rest of the session is allowed to do. The agent can keep reading files and answering the developer, and it can't run commands, write files, or make network requests until a person has looked at what happened.
How does session tainting work?
Session tainting has three parts:
- A finding sets a flag. Something screens the session's activity, such as the text a tool returned, and sets a named flag on the session when it finds a risk.
- Policies read the flags. Each later tool call carries the session's active flags, and a policy denies risky kinds of call while a flag is set.
- A person clears the flag. The session stays locked until an administrator reviews it and removes the taint. The agent can't clear it itself.
In Arcjet, the flags are set by Arcjet, never by the agent, and they arrive on each tool call as the session_flags policy input. The following flags can taint a session:
| Flag | Set when |
|---|---|
prompt_injection | A tool result looks like an instruction override or jailbreak |
sensitive_data_exposure | A tool result contains an entity type the policy denies, such as a card number or Social Security number |
credential_access | Session activity touched secrets, keys, tokens, or credential paths |
data_exfiltration | Local files or secrets were sent to an external host. On shell tools, this also requires upload evidence in the command, so an ordinary package install doesn't lock the session |
code_execution_risk | Untrusted code was fetched and run, piped into a shell, or used to escalate privileges |
elevated_risk | The session as a whole crossed a high risk threshold |
The first two flags come from screening tool results with Arcjet's prompt injection and sensitive information detectors. The rest come from Arcjet's risk scoring of the session's sequence of actions. Claude Code, GitHub Copilot, Cursor, and OpenAI Codex all report tool results through a post-tool-use hook, which is where the screening runs.
What does a session taint policy look like?
The coding-agent.tainted-session starter policy in Arcjet locks shell, web, MCP, and file-write calls while any of the flags is active, and still allows reads:
risky := {"shell", "web", "mcp", "file_write"}
lock_flags := { "prompt_injection", "credential_access", "data_exfiltration", "code_execution_risk", "sensitive_data_exposure", "elevated_risk",}
deny contains "tainted-session-risky" if { some flag in input.values.session_flags flag in lock_flags input.values.tool_kind in risky}You tune the policy by editing the two sets. For example, to lock only network access after an exfiltration finding, narrow risky to {"web", "mcp"} and lock_flags to {"data_exfiltration"}. The same policy declares which tool-result screens run: its prompt injection detector enables the prompt_injection flag, and its sensitive information detector, with the entity types to deny, enables sensitive_data_exposure.
How is a tainted session cleared?
In Arcjet, only a person in the Arcjet Console can clear a taint, and clearing it requires a written reason. The clear is recorded in the audit log, so every unlocked session has a named reviewer and an explanation. The agent can't clear its own taint through the hook or through MCP, and a SessionEnd event doesn't clear it either.
That review step is the point of the design. The finding that set the flag might be a false positive, in which case the reviewer clears it and the developer carries on. Or it might be the start of an incident, and the session's recorded activity shows exactly what the agent read and tried to do next.
What doesn't session tainting do?
- It doesn't undo the action that triggered it. A tool result has already reached the model when it's screened. Tainting stops what comes next; it can't withhold what the agent has read.
- It isn't set by a denied tool call. A policy that blocks the AWS CLI doesn't taint the session on its own. Flags come from tool-result screening and session risk scoring. To stop a specific action, use a per-call policy. For an example, see stopping coding agents using AWS credentials.
- It doesn't count events. A rule such as "taint after 10 commands in auto mode" isn't something a taint policy expresses, because flags come from findings, not from counters. To control auto mode, use a policy on the permission mode. For more information, see requiring human approval in Claude Code auto mode.
- It only runs the screens you publish. Tool-result screening runs only when a published policy reads
session_flagsand declares the matching detector.
How does session tainting compare with other controls?
| Control | Decides based on | Good at |
|---|---|---|
| Per-call policy | The current tool call's command, paths, and destinations | Stopping a known-bad action, such as reading SSH keys |
| Human approval | A person reviewing each risky action | High-stakes actions where a person can judge intent |
| Session tainting | Findings earlier in the same session | Stopping follow-up actions after untrusted content enters the context |
| Kill switch | An operator decision to stop an agent entirely | Incident response once a problem is confirmed |
The controls work together. Per-call policies stop the actions you can name in advance, session tainting contains the ones that follow an attack you couldn't prevent, and review turns each locked session into a decision a person made. For how these layers fit a coding agent rollout, see coding agent security tools compared.
Frequently asked questions
What is session tainting?
Session tainting marks an AI agent's session as untrusted after a risky finding, such as a prompt injection in a tool result, and then restricts later actions in that session – typically shell commands, web requests, MCP calls, and file writes – until a person reviews the session and clears the taint.
Why do AI agents need session tainting?
Per-call policies judge each action on its own inputs, so they can't see what the agent has already read. Once an injected web page or tool result is in the model's context it can't be withdrawn, and the next command may look ordinary. Tainting changes what the rest of the session is allowed to do.
What triggers a session taint in Arcjet?
Arcjet sets flags such as prompt_injection and sensitive_data_exposure by screening tool results, and credential_access, data_exfiltration, code_execution_risk, and elevated_risk from risk scoring of the session's actions. The agent never sets flags itself.
How is a tainted session cleared?
In Arcjet, only a person in the Arcjet Console can clear a taint, with a required reason that is recorded in the audit log. The agent can't clear it through the hook or MCP, and a SessionEnd event doesn't clear it.
What doesn't session tainting do?
It doesn't undo the tool result that triggered it, it isn't set by a denied tool call on its own, it doesn't count events, and it only runs the tool-result screens that a published policy declares.
AI runtime security in your code
Protect your AI agent workflows with Arcjet
Publish the tainted-session starter policy in dry run and see which follow-up calls it would lock.