Back to Blog
AI Agents

AI Agent Guardrails: Human-in-the-Loop, Permissions and Failure Handling

By Keyved Engineering Team··6 min read

Short answer

AI agent guardrails are the controls that limit what an agent can do and catch mistakes before they cause harm. A production guardrail stack classifies actions by risk, gives the agent least-privilege tools, validates every input and output, requires human approval above defined thresholds, enforces limits on steps, cost and time, fails safely with clear escalation, and logs every decision for audit.

Key takeaways

  • Classify every agent action by risk and apply controls proportionally.
  • Give agents narrow, purpose-built tools with the user's permissions — never admin access.
  • Put human approval on irreversible, financial or external-facing actions.
  • Design the failure path as carefully as the success path.

The difference between a chatbot and an agent is that an agent can do things: send emails, issue refunds, update records, move money. That's where the value is. It's also where the risk is.

An agent that drafts a wrong answer creates a bad reply. An agent that executes a wrong action creates a bad outcome: a refund to the wrong customer, a deleted record, a confidential document emailed to a supplier.

Guardrails are how you get the value without the risk. This guide covers the guardrail stack we put around every production agent.

What are AI agent guardrails?

Guardrails are the controls that limit what an agent can do and catch mistakes before they matter. They operate at several layers:

  1. Permissions — what the agent can access at all
  2. Validation — whether a specific input, output or action is allowed
  3. Human oversight — when a person must approve
  4. Limits — how much the agent can do in one task
  5. Failure handling — what happens when something goes wrong
  6. Audit — a record of everything the agent did and why

One rule underpins all of them: guardrails live in code, not in the prompt. Telling a model "never refund more than $500" is a suggestion. A function that rejects refunds over $500 is a guardrail.

How do you classify agent actions by risk?

Start by listing every action the agent can take and assigning a risk tier:

TierExamplesDefault control
ReadSearch knowledge base, look up order statusAllowed, logged, scoped to user's permissions
DraftPrepare an email, propose a journal entryAllowed; output shown to a human
Reversible writeAdd a note, tag a ticket, update a statusAllowed with validation, logged
External or sensitiveEmail a customer, share a document externallyApproval required until proven reliable
Irreversible or financialRefund, payment, deletion, contract changeAlways approval, or strict rule-based thresholds

This table becomes the backbone of your design. It also gives risk and compliance reviewers something concrete to sign off, which speeds approval.

How should AI agent permissions work?

Least-privilege tools

Give the agent narrow, purpose-built tools instead of generic powerful ones.

  • Bad: execute_sql(query) with database access
  • Good: get_order_status(order_id) and update_shipping_address(order_id, address) with validation

Narrow tools are easier for the model to use correctly and impossible to misuse beyond their scope. This applies equally to tools exposed through MCP servers.

Act as the user, not as an admin

When an agent acts for a user, it should inherit that user's permissions. A support agent working for a junior employee shouldn't see data that employee can't see. Shared service accounts with broad access are one of the most common — and dangerous — shortcuts.

Separate read and write

Keep read-only tools and data-changing tools separate, and grant write tools only to agents and workflows that need them.

How do you validate agent inputs and outputs?

Validate tool calls before executing

Every tool call should pass checks in code:

  • Schema: correct types and required fields
  • Business rules: amount limits, allowed values, valid state transitions
  • Ownership: the record belongs to the customer or user in context
  • Consistency: the action matches the conversation (a refund for an order that was actually discussed)

If validation fails, return a clear error to the agent so it can correct itself — or escalate.

Validate model outputs

Structured outputs should be parsed and validated against a schema. Free-text outputs going to customers may need checks for policy compliance, sensitive data and tone. Our LLM security guide covers improper output handling, one of the OWASP top risks.

Treat external content as untrusted

Emails, documents, web pages and tool results can contain hidden instructions. Keep them clearly separated from your instructions and never let them grant new permissions.

When should a human be in the loop?

Human-in-the-loop (HITL) means a person reviews or approves at defined points. Put it where the cost of a mistake is high:

  • Irreversible actions: payments, refunds, deletions
  • External communication: messages to customers, suppliers or regulators — at least initially
  • Low confidence: when the agent's own confidence or a validation check falls below a threshold
  • Novel situations: case types the agent hasn't handled before
  • Regulated decisions: credit, employment, healthcare — where people must stay accountable

Make approvals fast

Human review only works if it's easy:

  • Show the proposed action, the evidence and the agent's reasoning on one screen
  • Offer one-click approve, edit or reject
  • Put approvals where reviewers already work: Slack, the helpdesk, email
  • Record decisions — they become valuable evaluation data

Reduce approvals over time

Start conservative. As the agent proves reliable for specific case types, measured against your evaluation set, move those case types from "approve" to "auto with audit." This gradual earned autonomy is how most successful agent rollouts work.

What limits should every agent have?

Agents can loop, retry endlessly or wander off-task. Hard limits prevent runaway behaviour:

LimitWhy
Max steps or tool calls per taskStops loops
Max tokens or spend per taskStops cost blow-ups
Max wall-clock timeKeeps users from waiting forever
Rate limits per user and per toolContains abuse and runaway automation
Max value per action and per dayCaps financial exposure

When a limit is hit, the agent should stop and escalate, not try harder.

How should agents handle failure?

The failure path deserves as much design as the happy path.

  • Retries with backoff for transient errors such as timeouts and rate limits
  • Fallback models when a provider is down or throttled
  • Idempotent actions so retries never double-charge or double-send
  • Checkpointed state so a failed multi-step task resumes instead of restarting (multi-agent systems in production)
  • Graceful escalation: hand the case to a human with full context and a clear note of what failed
  • Honest messaging: tell the user what happened and what comes next

What should you log for audit?

For every task, record:

  • Who the agent acted for, and when
  • The inputs it received
  • Each model call, tool call, argument and result
  • Validation outcomes and any approvals, with the approver
  • The final output or action
  • Model and prompt versions

This is what lets you debug incidents, answer customer complaints, and meet regulatory expectations such as the record-keeping and human-oversight requirements for high-risk systems under the EU AI Act. It also feeds your observability stack.

A guardrails checklist for production agents

  • Every action classified into a risk tier
  • Narrow, purpose-built tools; no generic admin tools
  • Agent acts with the user's permissions
  • Tool calls validated in code against schema and business rules
  • External content treated as untrusted
  • Human approval on irreversible, financial and external actions
  • Limits on steps, spend, time, rate and value
  • Retries, fallbacks, idempotency and checkpointing
  • Clear escalation path with full context
  • Full audit log with model and prompt versions

How we build guardrails at Keyved

Guardrails aren't an add-on in our builds; they're part of the architecture from the first week. We map every action to a risk tier with your team, implement rules in code, and design approval flows that reviewers actually like using.

Our platform foundation provides the shared pieces — audit logging, rate and cost limits, retries and provider failover — so each agent project starts with them in place. See our AI agents and automation service and examples on our projects page.

Planning an agent that will take real actions? Talk to us about the guardrails before you build — it's much cheaper than adding them after an incident.

Frequently asked questions

What are guardrails in AI agents?

Guardrails are technical and process controls that constrain what an AI agent can do and detect mistakes. They include permission limits, input and output validation, content filters, human approval steps, rate and cost limits, and audit logging.

What is human-in-the-loop AI?

Human-in-the-loop AI means a person reviews, approves or corrects the AI's work at defined points before it takes effect. In agents, this is usually applied to high-risk actions such as payments, external communications or changes to important records.

How do you stop an AI agent from doing something wrong?

Limit the tools it can use and the data it can reach, validate every tool call against rules in code, require approval for risky actions, cap steps and spending, and monitor its behaviour so problems are caught early. Never rely on instructions in the prompt alone.

Do guardrails make AI agents slower?

Validation and rule checks add very little latency. Human approvals add waiting time, which is why they should be targeted at genuinely risky actions and reduced over time as the agent proves reliable for specific case types.

Which guardrail tools are available?

Options include schema validation libraries, provider moderation endpoints, open-source frameworks such as NVIDIA NeMo Guardrails and Guardrails AI, and policy engines. Most production systems combine these with custom business rules in code.

Want to see how we build these systems for clients?

Let's Talk

Keep reading