Arcjet vs OpenAI Agents guardrails

OpenAI Agents guardrails are SDK hooks where you write a check and signal a tripwire.

10 min read
In short: OpenAI Agents guardrails are SDK hooks where you write a check and signal a tripwire. Arcjet supplies the check: detectors, rate limits, and centrally managed policy, enforced before the run and on each authored function tool, with every decision recorded. Neither covers hosted tools or handoffs.

What is the difference between OpenAI Agents guardrails and Arcjet?

OpenAI Agents guardrails are hooks in the OpenAI Agents SDK where you write your own check and signal a result with a tripwire. Arcjet is a runtime security platform that supplies the check itself: centrally managed policy, rate limits, prompt injection and sensitive-information detection, and a recorded decision for every call. The key difference is who writes and changes the check: a developer, in agent code, for OpenAI guardrails, and a security team, from the Arcjet Console, for Arcjet. On OpenAI Agents, you call Arcjet with a direct guard() before the run and wrap each authored function tool with its adapter, and neither control covers hosted tools or handoffs. Choose OpenAI guardrails for SDK-native content checks that your team writes and maintains, and add Arcjet when a security team needs to set policy on tool calls without editing agent code and keep evidence of each decision.

Arcjet publishes this comparison. Competitor details come from OpenAI's public documentation for the OpenAI Agents SDK for JavaScript and for Python, reviewed on September 25, 2026.

OpenAI Agents guardrails and Arcjet are easy to confuse because both use the word guardrail. A team that adds an inputGuardrails check and then asks why refund_order still ran has mixed two jobs: the check screened the inbound text, and nothing authorized the tool call. The following sections describe where each control sits. For the implementation steps, see How do I secure an OpenAI Agents SDK agent?.

What do OpenAI input and output guardrails stop?

The OpenAI Agents SDK has three guardrail families, according to its guardrails guide:

  • Input guardrails (inputGuardrails) run on the initial user input, and only for the first agent in a workflow.
  • Output guardrails (outputGuardrails) run on the final output, and only for the agent that produces it.
  • Tool guardrails (defineToolInputGuardrail and defineToolOutputGuardrail in JavaScript, and the tool_input_guardrail and tool_output_guardrail decorators in Python) run before and after each function-tool invocation.

An input or output guardrail returns tripwireTriggered, and the SDK then throws an error such as InputGuardrailTripwireTriggered and halts the run. A tool guardrail returns one of three behaviors: allow, rejectContent to skip the call or replace its output with a message, or throwException to throw a tripwire error. The SDK also exposes callModelInputFilter as a filter on model input.

Two details from OpenAI's documentation affect what a guardrail can stop. Input guardrails run in parallel with the agent by default (runInParallel: true), and the guide notes that "the model may already have consumed tokens or run tools if the guardrail later triggers". Set runInParallel: false to block the model until the check completes. Tool guardrails apply to function tools that you define with tool() and to tools converted from local MCP servers that configure them, but the guide states that handoffs, hosted MCP tools, other hosted tools, and the built-in computerTool, shellTool, and applyPatchTool don't use that pipeline.

An OpenAI Agents guardrail is code that you write. A well-written input guardrail can refuse a jailbreak string before the model sees it, often by running a second agent as the classifier. The guardrail doesn't come with a detector, a shared rate-limit store, a policy that someone outside the repository can change, or a record of the decision.

What does an Arcjet in-process deny stop?

An Arcjet in-process deny stops the side effect itself, because Arcjet puts a decision in the path of the action, in your process, before the side effect runs. On OpenAI Agents, Arcjet decides at two points: before the run, and on each authored function tool.

A direct guard() call before the run screens user text that you already hold, such as with the prompt injection rule. On DENY, don't start the run. A direct guard() call fails open, so an ALLOW isn't proof that the rules ran. Check decision.hasFailedOpen() at a call site that must fail closed.

The Arcjet adapter wraps each authored function tool. In JavaScript, guardTool from @arcjet/guard/openai-agents/v0 wraps FunctionTool.invoke, so on DENY the original execute never runs. The JavaScript adapter returns a plain ArcjetDenialResult rather than throwing, and the runner records that result as the tool's output.

In Python, guard_tool from arcjet.guard.openai_agents attaches itself as a tool_input_guardrails entry and denies through reject_content(...), so Arcjet runs inside OpenAI's own tool-guardrail pipeline. Both adapters default to deny when Guard can't be evaluated.

Each wrapped call can carry an actor from your authenticated session, typed policy inputs, rate limits keyed on values you trust, and the SDK-local sensitive-information detector, which keeps the raw text in your process. A security team selects and changes the policy for each tool by its action label from the Arcjet Console or MCP server, and every decision is recorded. For the full contract, see the OpenAI Agents agent guard documentation.

The Arcjet OpenAI Agents adapter has documented limits. Hosted tools, MCP servers, handoffs, agents used as tools, and computer and shell tools don't go through the authored path, so they aren't Arcjet deny points, and neither Realtime nor Sandbox is covered. Runner tool-start and tool-end events are observe-only. There is no inbound helper and no approval helper. protect() remains the check for HTTP routes, and bot detection runs only there.

OpenAI output guardrails and the Arcjet adapter also interact. Because the JavaScript guardTool returns the denial as the tool's output, an outputGuardrails check and a customDataExtractor both receive the ArcjetDenialResult object, with arcjetDenied: true. Account for that before you write an output guardrail that assumes every tool result is real data.

Which questions separate OpenAI Agents guardrails from Arcjet?

OpenAI Agents guardrails and Arcjet differ on eight questions, from what each control is to how it behaves during an outage. The following table compares the two on each question.

QuestionOpenAI Agents guardrailsArcjet on OpenAI Agents
What it is

SDK hooks for checks that you write, signaled with tripwireTriggered or rejectContent

A runtime security platform: detectors, rate limits, and centrally managed policy, called in your process

Inbound text

inputGuardrails on the first agent, and callModelInputFilter

A direct guard() call before the run, with the prompt injection rule

Authored tool denyTool input and output guardrails on function tools

guardTool on FunctionTool.invoke in JavaScript, guard_tool on tool_input_guardrails in Python

MCP toolsTool guardrails on local MCP servers that configure themNot an Arcjet deny point on this adapter
Hosted tools and handoffsOutside the tool-guardrail pipeline, according to OpenAINot an Arcjet deny point on this adapter
Who changes the policyA developer, in the agent's code

A security team, from the Arcjet Console or MCP server, with dry run before live

EvidenceGuardrail results in the run and in tracing

Every decision recorded in the Arcjet Console, with SIEM export on the Enterprise plan

Human hold

needsApproval and hosted MCP requireApproval

Not a policy. See

needsApproval is not a policy

Outage behaviorDefined by your guardrail code

Direct guard() fails open. The tool adapters default to deny

For the inbound recipe, see How do I screen inbound prompts in OpenAI Agents SDK?. For the comparison between the two first-party SDKs, see OpenAI Agents SDK vs Claude Agent SDK.

When do you choose OpenAI guardrails, Arcjet, or both?

Choose OpenAI guardrails when you want content checks that stay inside the OpenAI programming model and your team is willing to write and maintain them. OpenAI guardrails are the only one of the two that can attach to a local MCP server's tools on this SDK, and an output guardrail can screen the final answer before it reaches your user.

Choose Arcjet when the decision on a tool call must reflect who the user is and what they are allowed to do, when a security team needs to change that policy without a deploy, or when you need a record of every decision. Arcjet also brings the same policy to other frameworks, to HTTP routes through protect(), and to the coding agents your developers run, including Claude Code, GitHub Copilot, Cursor, OpenAI Codex, and Muse Code, through their hooks.

You can run OpenAI guardrails and Arcjet on the same agent. A tripwire that rejects a jailbreak and an Arcjet deny on refund_order answer different questions. Neither covers hosted tools or handoffs, so keep resource-level authorization in the services those tools call.

What are the alternatives to OpenAI Agents guardrails?

Alternatives to writing your own OpenAI Agents guardrails include Arcjet, OpenAI Guardrails, NVIDIA NeMo Guardrails, and Guardrails AI. Each option covers a different part of the job:

  • Arcjet supplies the detectors, rate limits, and centrally managed policy, and enforces them in your process on authored tools, on HTTP routes, and in coding-agent hooks. Arcjet records every decision.
  • OpenAI Guardrails is a separate OpenAI library that validates inputs and outputs with configurable checks, including moderation, jailbreak detection, PII detection, and URL filtering, and offers a drop-in replacement for the OpenAI client.
  • NVIDIA NeMo Guardrails is an open-source toolkit with input, dialog, retrieval, execution, and output rails, including execution rails before and after an action.
  • Guardrails AI is an open-source Python framework that composes validators from its hub into guards on model input and output.

For a wider view of the market, see AI agent security platforms and the series hub, agent framework security.

Frequently asked questions

What is the difference between Arcjet and OpenAI Agents guardrails?

OpenAI Agents guardrails are hooks in the OpenAI Agents SDK where you write your own check and signal a tripwire. Arcjet supplies the check itself: detectors, rate limits, and centrally managed policy, enforced before the run and on each authored function tool, with every decision recorded.

Is Arcjet an alternative to OpenAI Agents guardrails?

Yes, for the checks that you would otherwise write inside a guardrail, such as prompt injection screening, sensitive-information detection, and per-user tool authorization. In Python, the Arcjet adapter attaches as a tool input guardrail, so it runs inside the SDK's own pipeline. Output guardrails on the final answer stay an OpenAI feature.

Can you use OpenAI Agents guardrails and Arcjet together?

Yes. A tripwire that rejects a jailbreak and an Arcjet deny on a refund tool answer different questions, and they run on the same agent. Neither covers hosted tools or handoffs, so keep resource-level authorization in the services those tools call.

Can Arcjet deny hosted tools or MCP on OpenAI Agents?

No. The Arcjet adapter gates authored function tools only. Hosted tools, MCP servers, handoffs, agents used as tools, and computer and shell tools aren't Arcjet deny points, and Realtime and Sandbox are outside the adapter. OpenAI's own tool guardrails can attach to local MCP servers.

Does a direct guard() call fail closed?

No. A direct guard() call fails open, so an ALLOW isn't proof that the rules ran, and you check hasFailedOpen() wherever the inbound screen must fail closed. The Arcjet tool adapters behave differently: on OpenAI Agents, both the JavaScript and Python adapters default to deny when the check can't be evaluated.

AI runtime security in your code

Protect your AI agent workflows with Arcjet

Arcjet runs inside your application, where it can use runtime context to enforce agent actions and budgets, detect prompt injection, and protect sensitive information before a workflow acts.