Submit

AI Agent Guardrails Explained: Types, Tools and How They Work

AI agent guardrails are the enforceable checks that keep autonomous agents inside safe limits. How OpenAI, LangChain, NVIDIA and AWS build them, and when to use each type.

Written by AiAgentsListing Team

•3 min read
AI Agent Guardrails Explained: Types, Tools and How They Work

What are AI agent guardrails

An agent that can only talk is easy to reason about. One that can refund a customer, open a pull request, or call an internal API is not: the moment it acts, someone has to answer what it is allowed to do and how you find out if it went outside that. That is the job AI agent guardrails do, and every major agent framework now ships its own version of them.

AI agent guardrails are checks that run before, during or after an agent's turn to validate input, block unsafe output, or gate a tool call. They can be a fast classifier that stops a prompt before an expensive model runs, a regex that redacts a credit card number, or a policy engine that decides whether a refund tool is allowed to fire.

How AI agent guardrails work

Across frameworks, guardrails cluster around the same few points in an agent's execution:

  • Input guardrails run on what reaches the agent before it reasons: user prompts and retrieved documents. They catch prompt injection, jailbreak attempts and off-task requests.
  • Output guardrails run on what the agent produces before it reaches a user or another system. They catch leaked secrets, policy violations and unsafe content a downstream system might execute.
  • Tool or execution guardrails run around a specific tool call, checking arguments before the call and the result after it. This is where an agent's ability to act, not just talk, gets bounded.
  • Dialog or policy guardrails constrain the conversation itself: which topics are off limits, which flows an agent must follow, which actions need human approval.

A guardrail can be deterministic (a regex, a keyword list, a rate limit) or model-based (a smaller LLM or classifier judging intent). Deterministic checks are fast and cheap but miss anything that doesn't match a pattern; model-based checks catch more but cost latency and money. Most production setups use both, deterministic first, model-based as a second pass.

Examples: guardrails in the major frameworks

OpenAI Agents SDK

The OpenAI Agents SDK (v0.22.3) attaches guardrails to agents and tools rather than to a single global filter. Input guardrails run only for the first agent in a chain; output guardrails run only for the agent producing the final answer. Input guardrails support two execution modes: parallel, where the guardrail runs alongside the agent for lower latency, and blocking, where it must finish before the agent starts so a triggered tripwire never lets the expensive model run at all.

Tool guardrails are separate: they wrap a FunctionTool and run an input check before the tool executes and an output check after, including on local MCP tools when the server configures them. When any guardrail's tripwire fires, the SDK raises an exception (InputGuardrailTripwireTriggered, OutputGuardrailTripwireTriggered, or the tool-specific equivalents) and the run stops, with RunConfig.output_guardrail_blocked_message letting a developer set the placeholder text a rejected output gets replaced with.

LangChain

LangChain's guardrails run through its middleware system, intercepting execution before the agent starts, after it finishes, or around model and tool calls. It ships a built-in PIIMiddleware that detects email addresses, credit card numbers (Luhn-validated), IP addresses, MAC addresses and URLs, with four handling strategies: redact, mask, hash, and block, plus separate flags to apply checks to input, output, or tool results. LangChain also ships built-in human-in-the-loop middleware that pauses before sensitive operations like financial transfers or production data changes so a person can approve them.

NVIDIA NeMo Guardrails

NeMo Guardrails is an open-source Python library, 7,170 GitHub stars on the NVIDIA-NeMo/Guardrails repo, that defines five rail types by where they sit in the pipeline: input rails (before the LLM: content safety, jailbreak detection, topic control, PII masking), retrieval rails (filtering RAG documents and chunks), dialog rails (flow control across turns), execution rails (validating tool call inputs and outputs, NVIDIA's own label for this is

Share:

Subscribe to our newsletter

One email a week. New agents, MCP servers and skills, and what is actually getting traction.

Read next