How to audit and log AI agent activity for compliance

One structured event per agent tool call: who asked, what ran, what your controls decided, and when.

10 min read
In short: Record a structured event for every AI agent tool call, including denials: the authenticated user, the action and resource, the control's decision, the outcome, and the time, with sensitive values redacted first. Arcjet Guards record the decision half in the tool handler and join it to your outcome events by decision ID.

How do I audit and log AI agent activity for compliance?

Record one structured event for every tool call an AI agent attempts. The event says who asked, what the agent tried to do and to which resource, when it happened, and what your controls decided, with sensitive argument values redacted before anything is written. Tie the events for one task together with a correlation identifier, and keep them somewhere with a retention policy that matches your obligation.

Arcjet is one way to produce the decision half of that record. Each Arcjet Guard call runs in the tool handler before the side effect and records the action's label, the authenticated actor, the rules evaluated, and an allow, deny, or error conclusion with a reason. Your application logs add the business outcome, joined to the Arcjet decision by its identifier.

A transcript of the conversation isn't an audit log. It shows what the model said, not what your systems allowed. For the evidence auditors ask for once the log exists, see compliance evidence for AI agent activity.

Record who, what, and when for every tool call

Log the attempt, not only the success. A denied or failed call is often the most important record, because it shows the control working. The following fields answer the questions an investigator or auditor asks first.

QuestionFieldSource
Who asked?User ID and tenant IDThe authenticated session, never a model argument
Which agent acted?Agent ID, agent version, and modelYour deployment configuration
What did it try?Tool name, as a stable action label such as tools.export-reportThe tool handler
On what?Resource type and ID, plus classified argumentsThe validated arguments, after redaction
What was decided?Conclusion, rule, reason, and decision IDThe policy check that ran before the call
What happened?Outcome, such as succeeded, failed, or skippedThe downstream call's result
When?Timestamps for the decision and the outcomeServer clocks, not the model's account of time
As part of what?Correlation ID for the agent run or taskThe run that started the tool chain

Take identity from server state. If the model can supply userId as a tool argument, then the log records whatever the model claimed, and so does any policy that reads it.

Use structured logging, one event per decision and outcome

Write events as JSON with a fixed schema, not as interpolated strings. A fixed schema lets you query "every export by this user last quarter" without parsing prose, and it stops a model-generated value from breaking the log format or injecting a fake line.

Emit two events per tool call: one when the control decides and one when the action completes. Separating them preserves a record of calls that were allowed but then failed, and of calls that were denied and never reached the downstream service.

{
"event": "agent.tool_call.decided",
"at": "2026-09-29T14:03:11.482Z",
"run_id": "run_8f2c1a",
"user_id": "usr_4417",
"tenant_id": "tnt_22",
"agent_id": "support-agent@1.8.0",
"action": "tools.export-report",
"resource": { "type": "report", "id": "rpt_9031" },
"decision_id": "gdec_01k6...",
"conclusion": "DENY",
"reason": "RATE_LIMIT"
}

Keep free text out of fixed fields. A tool argument such as an email body or a search query belongs in a separate, classified field, and only after redaction.

Redact sensitive fields before they reach the log

An audit log that stores every argument verbatim becomes a second copy of your most sensitive data, with broader read access and a longer retention than the source system. Decide what the log may hold before you decide how long to keep it.

Use an allowlist rather than a denylist. Log identifiers, enumerated values, amounts, and status codes by default, and drop every other argument unless you've classified it. For the free-text fields you must keep, such as a message subject a reviewer needs, redact detected entities first:

import { redact } from "@arcjet/redact";
const [subject] = await redact(args.subject, {
entities: ["email", "phone-number", "credit-card"],
});
log.info({
event: "agent.tool_call.completed",
action: "tools.send-email",
subject,
});

@arcjet/redact runs in your process, so the unredacted value isn't sent to a third party to be cleaned. Don't call the returned unredact function on the logging path; restoring the value re-creates the exposure. Apply the same rule to traces and error reports, which often capture arguments automatically. For language-level patterns, see how to redact sensitive data from logs.

Keep one exception in mind. If a sensitive-information check denied the call, log that it was denied and which entity type triggered it, never the matched value.

How Arcjet's decision trail supports the audit log

Arcjet records the decision half of each event for you. A guard() call in the tool handler stores the label, actor, rules, conclusion, and reason, along with any metadata and correlationId you pass. A capture() call after the side effect records what your application did and joins it to the decision through decisionId:

import { launchArcjet, tokenBucket } from "@arcjet/guard";
const arcjet = launchArcjet({ key: process.env.ARCJET_KEY! });
const exportBudget = tokenBucket({
refillRate: 10,
intervalSeconds: 3600,
maxTokens: 10,
});
export async function exportReport(
args: { reportId: string },
ctx: { userId: string; tenantId: string; agentId: string; runId: string },
) {
const decision = await arcjet.guard({
label: "tools.export-report",
actor: ctx.userId,
correlationId: ctx.runId,
// Server-controlled values only; never model-authored text.
metadata: { tenant_id: ctx.tenantId, agent_id: ctx.agentId },
rules: [exportBudget({ key: ctx.userId, requested: 1 })],
});
if (decision.conclusion === "DENY" || decision.hasFailedOpen()) {
return { error: "Export denied by policy" };
}
const report = await reports.export(args.reportId, ctx.tenantId);
arcjet.capture({
action: "report.exported",
decisionId: decision.id,
correlationId: ctx.runId,
metadata: { report_id: args.reportId, rows: report.rowCount },
});
return report;
}

The two shapes describe the same record. decision.id is the gdec_ value to store as decision_id in your own log event, the Guard label is that event's action, and the capture() action names the outcome, like an agent.tool_call.completed event. Write the ID into both so you can join the Arcjet record to your log line.

The actor comes from the session, so the record names the real user. The correlationId groups every guard and capture call in one agent run, which is how you reconstruct a chain of tool calls after an incident. A dry-run rule is recorded too, so you can show that a threshold was measured against real traffic before it started to block.

To retrieve the trail, use the Arcjet Console's guard activity view, the CLI (arcjet guards list, arcjet guards details, and arcjet guards watch), or the MCP server's list-guards and get-guard tools. Decisions for HTTP routes protected with protect() sit alongside them under arcjet requests list and arcjet requests explain.

Map the audit trail to SOC 2 and audit-trail requirements

SOC 2 doesn't prescribe an AI agent log format, and your auditor decides what evidence satisfies each criterion. The agent audit log is commonly offered as evidence for the following Trust Services Criteria.

CriterionWhat it asks, in shortWhat the agent audit log shows
CC6.1 and CC6.3Logical access is restricted and granted by role

Each tool call was checked against the actor's permissions before it ran, including the denials

CC7.2System components are monitored for anomalies

Rate-limit, injection, and sensitive-information decisions per action and per identity

CC7.3 and CC7.4Security events are evaluated and responded to

A correlation ID that reconstructs the full tool chain for one agent run

CC8.1Changes are authorized and tested

Rules in reviewed code, and dry-run decisions recorded before enforcement

C1.1 and P-seriesConfidential and personal information is protected

Redacted arguments, and denial records that omit the matched value

The same fields cover the audit-trail sections of other frameworks, such as ISO/IEC 27001 Annex A 8.15 (logging) and the logging and traceability expectations in the EU AI Act for high-risk systems. Map them with your compliance owner rather than treating the table as a certification claim.

Arcjet itself has completed a SOC 2 Type 2 examination covering Security, Availability, and Confidentiality. The report is available through the Trust Center.

Keep the log for as long as the obligation requires

The Arcjet Console retains decision history according to your plan's log retention, which is longest on Enterprise. Treat the Console as the operational view, not the archive. On the Enterprise plan, Arcjet exports decisions to Datadog, Splunk, SentinelOne, Panther, and Amazon S3, where your own retention policy, alerting, and access controls apply.

Store the application events that carry the same decision ID in that system too, so the decision and the outcome join in one place. Restrict who can read the audit store, make it append-only where you can, and test that a record for a known tool call can be found by user, action, and run.

Checklist

  • Log every tool call attempt, including denials and errors.
  • Take user, tenant, and agent identity from server state.
  • Use a fixed JSON schema with a stable action label for each tool.
  • Emit a decision event and an outcome event, joined by a decision ID.
  • Allowlist logged fields, and redact free text before writing it.
  • Pass a correlation ID for each agent run.
  • Export to a store with a retention policy that matches your obligation.
  • Map the fields to your framework's criteria with your compliance owner.

AI agent governance guides

This guide is one of five on governing what an AI agent may do at runtime:

Learn more: Arcjet Guards · Security and compliance

Frequently asked questions

How do I audit and log AI agent activity for compliance?

Write one structured event for every tool call attempt, including denials. Record the authenticated user and tenant, the agent, the action and resource, the control's decision and reason, the outcome, and timestamps, joined by a correlation ID for the agent run. Redact sensitive values before writing, and export to a store with the retention your obligation requires.

What should an AI agent audit log contain?

Who asked (user and tenant from the session), which agent acted, what it tried (a stable action label), on which resource, what the control decided and why, what happened downstream, and when. Classified arguments are optional; raw free text is not.

How do I keep sensitive data out of AI agent logs?

Allowlist the fields you log, and redact detected entities in any free text you keep, in your own process, before it is written. If a sensitive-information check denied a call, log the entity type that triggered it, never the matched value.

Is an AI agent audit log enough for SOC 2?

It is common evidence for access-control, monitoring, and change-management criteria such as CC6.1, CC7.2, and CC8.1, but your auditor decides what satisfies each criterion. Map the fields with your compliance owner rather than treating the log as a certification.

How does Arcjet help audit AI agent activity?

Each Arcjet Guard call in a tool handler records the action label, the authenticated actor, the rules evaluated, and an allow, deny, or error conclusion with a reason, including dry-run decisions. capture() records what your application did and joins it to that decision by ID, and Enterprise plans export decisions to Datadog, Splunk, SentinelOne, Panther, and Amazon S3.

AI runtime security in your code

Protect your AI agent workflows with Arcjet

Capture events record what the application did after a decision, which is the half of the evidence most teams are missing.