Producing Compliance Evidence for AI Agent Activity

Auditors do not want your system prompt.

7 min read
In short: An auditor asks four things: that a control exists at the point of risk, that it was evaluated before enforcement, that it operates with reasons, and that someone owns it. What is usually missing is a record of the decisions your controls made.

What compliance evidence do you need for AI agent activity?

Auditors don't want your system prompt. They want evidence that a control existed at the point of risk, that it was evaluated before it was enforced, that it operates and produces decisions with reasons, and that someone owns it and changes it deliberately.

A record of what your agent did covers part of one of those. What is usually missing is a record of the decisions your controls made on each action. For how to produce that record, field by field, see how to audit and log AI agent activity for compliance.

What an auditor is actually asking for

"How do you govern your AI agents?" isn't a question about model behavior. It decomposes into the four asks in the following table.

The askWhat satisfies itWhat does not
A control exists at the point of risk

The check is in the code path of the consequential action, shipped in the same change as the tool

A policy document, or a prompt instructing the model to behave
It was evaluated before it was enforced

Dry-run decisions recorded against real traffic, before the rule started blocking

An assertion that the threshold was chosen carefully
It operates

Decisions with allow, deny, or error conclusions and reasons, retrievable per action and per identity

Application logs of what the agent did, which show outcomes but not the control's decisions

Someone owns it

Rules in reviewed code, thresholds changed deliberately through a named process

Configuration whose history nobody can account for

The evidence a runtime decision produces

Each Arcjet decision records an allow, deny, or error conclusion, along with the rule that fired and the reason. DRY_RUN is a rule mode, not a fourth conclusion: a dry-run denial is recorded without changing the conclusion to deny. Guards add a label naming the action, such as tools.issue-refund, and a metadata map that you populate with the context that matters. If you pass a correlationId, then Arcjet ties the decision to the workflow run that produced it. That identifier is for reconstruction; it doesn't change the allow or deny.

That maps onto the four asks directly. The control exists because the check is in the code path, in the same pull request as the tool. It was evaluated first because dry-run results are recorded too, which makes "we measured before enforcing" demonstrable rather than asserted. It operates because decisions carry reasons and are retrievable per action and per identity. Ownership holds because rules live in reviewed code, and Guard policies (a labeled-action system, plus coding-agent policies selected by moment) let a security team change expression rules and detectors without a deployment. Token-bucket stays in application code.

Keep metadata to server-controlled values: authenticated user and team identifiers, enumerated types, status codes, amounts. Never include user-authored free text. Downstream consumers treat metadata as untrusted, and it is the part of a record most likely to leak something.

Retrieve the evidence

Use the tools that match the enforcement path.

Console filters by time and conclusion. Request decisions search by host, path, or request ID. Guard decisions have their own activity view.

CLI. For protect() on HTTP routes: arcjet requests list, arcjet requests details --request-id <id>, and arcjet requests explain --request-id <id>. For guard() on tools, MCP, and jobs: arcjet guards list, arcjet guards details, and arcjet guards watch, which streams live guard activity. arcjet analyze dry-run-impact is the pre-enforcement picture for request remote rules.

MCP server. Request path: list-requests, get-request-details, explain-decision. Guard path: list-guards, get-guard. get-dry-run-impact remains the request-rule promotion check.

Be precise about the limits

Arcjet retains decision history in the Arcjet Console according to your plan's log retention (longest on Enterprise). For anything longer, or for detection and alerting, export the decisions to your SIEM or log store. On the Enterprise plan, Arcjet sends decisions to Datadog, Splunk, SentinelOne, Panther, and Amazon S3.

Don't assume the Console record covers a long retention obligation on its own. If your obligation requires seven years of retained evidence, the copy in your SIEM or S3 bucket, under your own retention policy, meets that obligation. The Console is the live operational view. Anyone describing the Console alone as an audit trail in the archival sense is overclaiming, and an auditor who probes it finds the gap.

Export gives you the durable copy with the decision identifier, rule, and reason on every record. Where you also act on a decision in your own code, emit your own event carrying the same decision identifier, so the two records join.

The platform half

Evidence about your controls helps only if the platform enforcing them is itself in scope.

Arcjet has completed a SOC 2 Type 2 examination covering Security, Availability, and Confidentiality with an unqualified opinion, with the report available through the Trust Center.

Because SDK-local sensitive-information inspection keeps the raw request body in your process, your assessment doesn't have to account for protected data leaving your environment for a third party to check it on that path. The server-side sensitive-information detector (coding-agent hooks and other no-SDK policies) is a different claim. That removes a transfer question from the analysis rather than answering it. For detail about which controls run locally and which call out, see keeping security inspection local.

Checklist

  • Put the control in the code path of each consequential action, not in a document.
  • Run in dry run first, so the pre-enforcement measurement is itself recorded.
  • Label every check with the action it protects, using a stable slug.
  • Populate metadata with server-controlled values only.
  • Tag decisions with the workflow run through a correlationId.
  • Export decisions to your SIEM or S3, and emit your own event carrying the decision identifier where your code acts on a decision.
  • Know your plan's Console retention window (longest on Enterprise), and keep the archival copy in the system you export to.
  • Keep threshold changes deliberate and attributable.

Where Arcjet fits

Compliance evidence for agents is split between two kinds of tooling. Governance and posture platforms such as Cranium and Refractal describe your AI estate and its controls. Observability platforms record what the agent did.

Arcjet contributes a narrower and often missing piece: the record of what your enforcement controls decided, per action, with the rule and reason, including dry-run decisions from before enforcement was switched on.

Console retention follows your plan (longest on Enterprise). On the Enterprise plan, decisions export to Datadog, Splunk, SentinelOne, Panther, and Amazon S3, which is where detection rules, alerting, and long-term retention belong. A SOC 2 Type 2 examination covers the platform itself, with the report in the Trust Center.

Learn more: Security and compliance · Architecture

Frequently asked questions

What compliance evidence do I need for AI agent activity?

Evidence that a control existed at the point of risk, that it was evaluated before it was enforced, that it operates and produces decisions with reasons, and that someone owns it and changes it deliberately. A record of what the agent did covers only part of the third.

How long is Arcjet decision history retained?

Console history follows your plan's log retention, which is longest on Enterprise. For longer retention, detection, and alerting, export decisions to Datadog, Splunk, SentinelOne, Panther, or Amazon S3 (Enterprise plan) and keep the archival copy there.

Can I send Arcjet decisions to my SIEM?

Yes. On the Enterprise plan, Arcjet exports decisions to Datadog, Splunk, SentinelOne, Panther, and Amazon S3, so you can build detection rules and alerts on agent activity in the tools your security team already uses.

Can I describe Arcjet decision records as an audit trail?

The Console view alone is not an archival audit trail; it covers your plan's retention window. Exported to your SIEM or S3 under your own retention policy, the same records, each carrying its decision identifier, rule, and reason, form the durable copy.

How do I retrieve Guard decisions versus request decisions?

Use Guard tools for tool, MCP, and job decisions: CLI arcjet guards list, details, and watch, and MCP list-guards and get-guard. Use request tools for HTTP request decisions: arcjet requests list and MCP list-requests, get-request-details, and explain-decision.

How do I demonstrate a control was measured before enforcing?

Run it in DRY_RUN mode first. DRY_RUN is a rule mode, not a third conclusion: dry-run denials are recorded without changing the allow or deny. That makes the pre-enforcement measurement demonstrable rather than asserted.

AI runtime security in your code

Protect your AI agent workflows with Arcjet

Capture events record what the application did after a decision, which is the half of the evidence most teams are missing.