AI agents Playbook

AI Agent Guardrails: 8 Controls That Hold Up in Production

Guardrails are the controls between what an agent decides and what it is allowed to do. Eight of them decide whether an agent can be trusted with live work, and every one of them lives in code, outside the model.

For CTOs and CISOs deciding what an AI agent may do on its own before it touches customers, money or production.

Published
Reviewed
Reading time
13 min

The short answer

AI agent guardrails are the controls that sit between what an agent decides and what it can do. The ones that hold in production are enforced in code, outside the model: scoped tool permissions, untrusted-input handling, action classes with approval for irreversible steps, idempotent writes, full traces, evaluation on every change, a designed exception path and autonomy that widens only as measured error falls.

Key takeaways

  • A prompt asks the model to behave. Code makes it. Every control that matters runs outside the model, where no input can argue its way past it.
  • Class every action as read, reversible or irreversible before the agent may take it. Irreversible steps wait for a named person.
  • Treat everything the agent reads as untrusted. OWASP ranks agent goal hijack, carried in through content, first among agentic risks.1
  • Autonomy is earned in stages: shadow, assist, bounded, then wider limits as the measured override rate falls.
  • Gartner expects over 40% of agentic AI projects to be canceled by the end of 2027, and names inadequate risk controls as one of three causes.2

A demo works because the happy path is short and the reviewer is generous. Production is the opposite. Inputs arrive malformed, an API times out halfway through a write, a customer asks for something the agent was never designed for, and the same request lands twice because someone clicked again. Guardrails are what turn an impressive loop into a system your risk team can sign off and your support team can defend.

  • 40%+of agentic AI projects Gartner expects to be canceled by the end of 20272
  • 97%of organizations with an AI-related breach lacked proper AI access controls3
  • 23%of respondents say their organization is scaling an agentic AI system somewhere in the enterprise4

What AI agent guardrails are

A guardrail is any control that limits what an agent can see, decide or do, and records what it did. The useful ones sit in three layers: what the agent may touch, what it may do without a person, and what the business can see and reverse afterwards. A system prompt belongs to none of those layers. It is a request the model usually honors, which is a fine way to shape tone and a poor way to protect a ledger.

Exhibit 1Where a control lives decides whether it holds

In the prompt

Instructions the model reads

  • Followed most of the time
  • Can be argued with by the content the agent reads
  • Changes behavior when the model version changes
  • Leaves no record of the decision it made

In code, outside the model

Checks every tool call must pass

  • Applied to every call, for every input
  • Cannot be talked out of by a document or an email
  • Survives a model swap unchanged
  • Logs the decision with the policy version that made it
Every guardrail in this playbook lives in the right-hand column.

1. Enforce tool permissions in code

Asking a model not to do something is a suggestion. The enforcement has to live in the layer that executes tools. Every tool call is checked against the identity the agent is acting for, exactly as it would be in your application: an agent working on behalf of a support representative gets that representative's authority and no more. Permission to call a tool says nothing about its arguments, so recipients, amounts and record ids are validated against a schema and checked against limits. OWASP names excessive functionality, excessive permissions and excessive autonomy as the three roots of agents doing damage, and all three are fixed here.5

// The policy layer runs on every call. It is not optional, and it is not in the prompt.
async function invoke(tool: Tool, input: unknown, actor: Actor) {
  assertAllowed(actor, tool.permission);        // authorization, as for a person
  const args = tool.schema.parse(input);        // shape, limits, allowed recipients
  const cls = classify(tool, args);             // read | reversible | irreversible
  if (cls === "irreversible") return queueForApproval(tool, args, actor);
  return tool.run(args, { idempotencyKey: keyFor(tool, args, actor) });
}
A sketch of the call path: authorization, validation and classification happen before any tool runs.

Permissions decide what the agent may do. Limits decide how much of it, and they bound the damage from any failure the other controls miss. Each one is enforced in the same layer and set by the owner of the process:

LimitWhat it containsWhat happens at the limit
Steps per taskA loop that never converges on an answerThe task stops and goes to a person with its trace
Spend per task and per dayRunaway model and API costAn alert at a set share of the cap, a stop at the cap
Records touched per runOne bad instruction applied to a whole tableA small batch runs; the rest waits for review
Value per action and per dayMoney or stock moved without a personAnything above the limit becomes irreversible and waits for approval
Calls per tool per minuteAn agent overwhelming a downstream systemThe same throttle you apply to any other client of that system
Limits live in the policy file beside the permissions, versioned and reviewed like code.

2. Treat every input as untrusted

Two things in an agent loop carry authority: your system instructions and the request from the person the agent acts for. Everything else it reads is data, including retrieved documents, tool responses, email bodies, web pages and API payloads. Label that material with where it came from and keep it apart from the instructions. Text in a retrieved document or a customer email can ask for anything. The permission layer refuses it, which is why that boundary belongs in the architecture.

OWASP's Top 10 for Agentic Applications ranks agent goal hijack first: an attacker redirects the agent through content it reads, and the agent pursues the attacker's goal while it believes it is serving the user.1 No prompt fully prevents it. What limits the damage is the first guardrail: an agent hijacked by an email still cannot call a tool it was never given.

3. Class every action by consequence

ClassExamplesWhat happensWho approves
ReadLook up an order, search the policy library, query the warehouseRuns on its own, loggedNobody
ReversibleUpdate a CRM field, draft a reply, open a pull request, isolate a laptopRuns on its own, with a one-step undoNobody at the time; sampled in review
IrreversibleMove money, deploy to production, message customers, delete recordsWaits in the approval queueA named role, within limits the business sets
The three action classes. The threshold for each tool is a business decision, written into a policy file and versioned like code.

Most of the design work is deciding, with the business, which class each action belongs in and where the limits sit. A refund under a set amount may be reversible in practice because it can be clawed back; one above it is not. Those are commercial decisions, and they belong in a policy file your risk owners review and approve, never in the model's judgment on the day.

An approval gate protects the business only if the approver reads what they approve. A queue that shows a raw tool call gets approved on reflex. Each request should state the action in business terms, the evidence the agent relied on, the policy line that required approval and the undo, if one exists. Separate duties, so the person who gave the agent the task cannot approve its irreversible step above a limit. Expire approvals that wait too long, so a stale decision never runs against data that has since changed.

4. Make every write idempotent

Agents retry, and networks fail mid-call. Derive the idempotency key from the intent, so every retry of the same refund carries the same key and the payment system processes it once. Without it, a timeout followed by a retry issues a second refund. This is among the most common production incidents in agent systems, and it is entirely preventable at design time.

5. Trace every step

An agent's decisions are only as defensible as the record of them. Every task should leave a trace that a reviewer can replay months later:

  • The inputs the agent saw, including retrieved context and where it came from.
  • The tools it considered, the one it chose and the arguments it passed.
  • The result each tool returned.
  • The identity it acted for, and the policy version that allowed each action.
  • Who approved each irreversible step, and when.
  • A correlation id that ties the task together across services, in the open OpenTelemetry format your dashboards already read.

6. Evaluate on every change

Build a graded set from your real historical cases, including the awkward ones your team argued about. Run it in CI on every prompt change, model change and integration change, and fail the release when the pass rate drops. When a better model is released, you learn in an afternoon whether it helps your cases or hurts them, instead of learning it from complaints.

7. Design the exception path first

Decide before launch what happens when confidence is low, a tool fails or a request falls outside scope. The agent stops, hands the case to a person with the context already assembled, and never guesses at an action with consequences. A system that fails into a human queue is one you can trust with your ledger; a system that fails silently is one you will switch off.

8. Ramp autonomy in stages

Exhibit 2Autonomy is earned, one stage at a time
  1. ShadowThe agent runs on live traffic. Its decisions are recorded and compared with your team's, and none is executed.Moves up when agreement with your team is high on your graded cases.
  2. AssistThe agent prepares each action; a person approves it with one click.Moves up when the override rate is low and falling.
  3. BoundedReads and reversible actions run on their own, within limits. Irreversible ones still wait.Limits widen as the measured error rate stays under target.
  4. WidenedLimits rise where the record supports it. One switch returns the agent to assist mode.Tested before it is needed.
Each stage runs on live work, and each step up is a decision made on measured numbers, approved by the owner of the process.

How to tell whether your guardrails are working

Measure guardrails the way you measure any other control: by what they catch, what they cost the operation and how quickly they let you answer a question about a decision. Each of these comes straight from the traces, so it can sit on the same dashboard as the agent's throughput.

MeasureWhat it tells youLook into it when
Blocked tool callsHow often the agent reaches past its permissionsIt rises after a prompt, model or tool change
Override rate in assist modeHow often your team corrects the agentIt stops falling, or starts to rise
Approval rejection rateWhether approvers are reading what they approveIt sits near zero on a busy queue
Time in the approval queueWhat the gate costs the operationIt grows faster than volume
Hand-off rateHow often the agent routes a case to a personIt jumps, which usually means the inputs changed
Time to replay a decisionWhether the traces are completeAnswering any question takes more than a query
Six measures of the controls themselves, alongside the agent's own quality and cost metrics.

Where AI agent guardrails fail in practice

Guardrails usually fail through a shortcut taken under deadline, and the same six shortcuts account for most of the gaps we find when we review an agent built elsewhere.

ShortcutHow it shows upThe fix
The rule lives in the promptThe agent follows it until a document tells it otherwiseMove the rule into the call path, where content cannot reach it
One shared service accountEvery action logs as the same identity, so nobody can say whom the agent acted forAct on behalf of the user, with that user's permissions
Approval on reflexApprovals take seconds and are almost never rejectedShow business context, separate duties and sample approved actions
Logs without replayThe trace records the answer and loses the inputsLog inputs, tool results and policy version under one correlation id
A graded set that never changesThe pass rate stays high while complaints riseAdd every production failure to the set as a new case
A stop switch nobody has pressedThe switch exists on the architecture diagramExercise it on a schedule, the way you test a failover
Each shortcut saves days before launch and costs far more in the first incident.

Guardrails limit what an agent can do once something goes wrong. The threat model that explains how it goes wrong, from injected instructions to poisoned tools, is in AI agent security, and building the graded set behind control six is covered in LLM and AI agent evaluation.

Map the controls to the frameworks your reviewers use

Security and compliance reviewers will ask how the agent maps to the standards they already work from. The eight controls line up cleanly with the OWASP agentic list, the NIST AI Risk Management Framework and, for systems in scope, the EU AI Act.178

ControlOWASP agentic risks addressedNIST AI RMFEU AI Act
Tool permissions in codeTool misuse; identity and privilege abuseManageArticle 15, robustness and cybersecurity
Untrusted input handlingAgent goal hijack; memory and context poisoningMeasure, ManageArticle 15
Action classes and approvalsHuman-agent trust exploitation; rogue agentsGovern, ManageArticle 14, human oversight
Idempotent writesCascading failuresManageArticle 15
TracesAll ten, as evidenceMeasureArticle 12, record-keeping
Evaluation on every changeRogue agents; cascading failuresMeasureArticles 9 and 15
Exception path and staged autonomyCascading failures; rogue agentsGovern, ManageArticles 9 and 14
Our mapping of the eight controls to the three frameworks. Article numbers refer to the obligations for high-risk AI systems.

The EU deadlines moved in 2026. Regulation (EU) 2026/1744 pushed the high-risk obligations for Annex III systems to December 2, 2027 and for AI inside regulated products to August 2, 2028.9 The records those obligations require are the same traces and approvals described above, and they are far cheaper to build in than to reconstruct later.

What to ask before an agent goes live

Six questions for the go-live review

  • Which tools can the agent call, and where is that list enforced?
  • Which actions are irreversible, and which named role approves each one?
  • What does the agent do with a request it does not understand?
  • Can we replay any decision it made last month, with the inputs it saw?
  • What is the pass rate on our graded cases, and what fails a release?
  • How quickly can one person return the agent to assist mode?

None of this is exotic. It is the discipline you already apply to any system that touches money or customers, applied to a component that makes its decisions probabilistically. Teams that skip it are borrowing time from their own incident response; teams that build it in are the ones whose agents are still running a year later.

We build these eight controls into every agent we put into production as part of our AI agent development work. The policy file, the approval queue, the traces and the graded set are delivered with the agent, and your risk owners review each of them before it goes live.

Questions leaders ask

What is the difference between AI guardrails and AI governance?

Governance is the set of decisions a company makes about AI: who owns each system, which risks are acceptable and who approves what. Guardrails are how those decisions are enforced inside a running system, as permissions, approval gates, limits and traces. Good governance names the approver for an irreversible action; the guardrail is the gate that holds the action until that person approves.

Can prompt injection be fully prevented?

No. Models cannot reliably separate instructions from the content they read, which is why OWASP ranks goal hijack first among agentic risks.1 The defense is architectural: keep untrusted content apart from instructions, give the agent only the tools the task needs, and hold irreversible actions for approval, so a hijacked agent has little it can do.

Which actions should always need human approval?

Any action that cannot be undone or that commits the business: moving money above a set limit, deploying to production, sending messages to customers or regulators, and deleting records. The exact limits are commercial decisions, recorded in a policy file and approved by the owner of the process.

How do you test AI agent guardrails before launch?

Test the policy layer like any security control. Unit tests call every tool with arguments outside its limits and confirm each call is refused. Replayed historical cases run through the whole loop. Adversarial text planted in documents and tool responses checks whether the agent attempts an action it should not. Score whether the harmful action completed, since a polite refusal means nothing if the call went through, and repeat the suite after every change.

Are content filters enough to act as AI agent guardrails?

Content filters are one layer. They inspect the text going into and out of the model, which catches harmful output and some injection attempts. They cannot see what a tool call will do to your systems. Permissions in the call path, limits, approval for irreversible steps and traces are what contain the consequences, and they keep working on the day a filter misses something.

Do guardrails make an AI agent slower?

Permission checks, validation and tracing add milliseconds to a tool call. Approval gates add time only to irreversible actions, which are usually a small share of the work. The larger effect runs the other way: agents with strong controls are allowed to do more on their own, sooner, because the business can see and reverse what they do.

Sources

  1. OWASP Top 10 for Agentic Applications for 2026OWASP GenAI Security Project, December 2025
  2. Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027Gartner, June 25, 2025
  3. IBM Report: 13% of Organizations Reported Breaches of AI Models or Applications, 97% of Which Reported Lacking Proper AI Access ControlsIBM, Cost of a Data Breach Report 2025, July 30, 2025
  4. The state of AI in 2025: Agents, innovation, and transformationMcKinsey & Company, November 5, 2025
  5. LLM06:2025 Excessive AgencyOWASP Top 10 for LLM Applications 2025
  6. BC Tribunal Confirms Companies Remain Liable for Information Provided by AI ChatbotAmerican Bar Association, Business Law Today, February 2024
  7. AI Risk Management FrameworkNational Institute of Standards and Technology
  8. Regulation (EU) 2024/1689, the Artificial Intelligence ActOfficial Journal of the European Union
  9. Regulation (EU) 2026/1744, the Digital Omnibus on AIOfficial Journal of the European Union, July 24, 2026

Written by DigyAi Engineering from the systems we build and run. Every figure links to its public source, and every link and figure was checked on September 26, 2026. No client data appears in our insights.

Read next

All insights
  • An AI agent on its own tower reads web pages, email and tool text from an island outside; every action it plans passes through a lit policy engine, which lets calls to the company's systems through and stops a call to send data out at a lowered barrier. A kill switch is wired to the agent at the front.

    Security, risk and compliance Guide

    AI Agent Security: The Threat Model and Controls a CISO Should Require

    For CISOs and security architects deciding whether an AI agent is safe to connect to production systems, company data and customers.

    16 min read

  • Your own cases, stored in a golden-set archive, run through your system to a release gate whose board shows every segment against a threshold its owner signed in advance; one segment falls short and the gate holds, then the rerun clears the line and the release crosses a bridge to production.

    LLM and RAG engineering Guide

    LLM and AI Agent Evaluation: How to Prove a System Is Ready to Ship

    For CTOs, heads of AI and risk owners deciding whether an LLM application, RAG system or AI agent is ready to leave the pilot, and what evidence should back that decision.

    16 min read

  • A chatbot on its own small island answers on a screen, and its only lane ends at a red stop at the island's edge; on the main plinth an AI agent tower cancels the order, queues the refund, notifies the customer and logs each step, every system ticked.

    AI agents Explainer

    AI Agents vs Chatbots vs Agentic AI: What Actually Differs

    For CTOs and business owners deciding whether a workflow needs an AI agent, a chatbot or plain automation before they fund the build.

    16 min read

Get in touch

Tell us what you are building.

Write it as big as you imagine it.