Point of view3 min readAI Design · Product & Experience Design · Technology & Intelligence

Designing agents people can trust: guardrails, approvals and a way back

An agent that can act is only as good as the limits around it. Seven design rules we use to keep agents useful, honest about what they do not know, and unable to do harm on their own.

Xterra Edze studioEditorial

Published

A winding mountain road with a guardrail along its edge.
A guardrail does not slow the road down. It is what makes the curve safe to take.Photo: James Coleman on Unsplash

An assistant that answers questions can embarrass you. An agent that sends emails, changes records or spends money can hurt someone. The design work that matters most is not the conversation. It is deciding what the agent may do on its own, what it must ask about, and how a person sees, stops and reverses what it did.

One job, the fewest tools

Scope the agent to one job and give it only the tools that job needs, with the narrowest permissions those tools allow. The OWASP Top 10 for LLM Applications names this risk directly as excessive agency, and traces it to three causes: excessive functionality, excessive permissions and excessive autonomy.1 An agent that drafts replies does not need to send them. One that reads a CRM does not need to delete from it.

Treat every input as untrusted

Prompt injection is first on the same list.2 Instructions can arrive in what the user types, and also in what the agent reads: a web page, a PDF, an email it summarises. Design so that no input can widen what the agent is allowed to do. Permissions live in the systems the agent calls, never only in its prompt.

Show the plan before it runs

For anything consequential, the agent proposes and a person approves. The approval screen shows what will happen, to which records and why, in plain words, with the sources the agent used. Approving should be one decision, not a hunt through a transcript.

Two men looking intently at a computer screen.
An approval is a screen designed for one decision: what will happen, to what, and on what evidence. Photo: litoon dev on Unsplash

Make uncertainty visible

Research on human–AI interaction has long asked systems to make clear what they can do and how well they can do it.3 4 In practice that means citing sources, saying when the answer is not in them, and handing the question to a person instead of guessing. A clean refusal is a feature.

Every action has a way back

  • Draft first: emails, records and posts are created as drafts until a person approves them.
  • Undo for a set period after any change the agent makes.
  • A dry run that shows the effect without making the change.
  • A stop control that ends the run and leaves everything as it was.

Log everything a reviewer would ask for

Inputs, retrieved sources, tool calls, approvals, refusals and errors, each with a time and an actor. NIST’s Generative AI Profile organises its guidance around four primary considerations — governance, content provenance, pre-deployment testing and incident disclosure — and none of them works without that record.5 The log is also how you learn where the agent helps and where it hesitates.

Test before anyone relies on it

Before release, run the agent against an eval set: real tasks with known good outcomes, plus the cases it must refuse and the attacks it must ignore. Release when it passes, and run the set again on every change to a model, a prompt or a tool.6 We wrote about building those sets in Evals before prompts.

Job per agent
1
With the fewest tools that job needs.
Irreversible actions without a person
0
Sending, paying, deleting.
Of steps written to the audit log
100%
Inputs, sources, tool calls, approvals.

The design rules we set for agents we build. Rules, not measured results.

The design work that matters most is not the conversation. It is deciding what the agent may do on its own.

We design agent experiences in AI Application Design and build them, with their evals and logs, in AI Strategy & Agents.

Dates. What this guide tracks.

The changes this article covers, in order, each with the source that sets the date. See every date in the standards ledger

  1. NIST publishes its Generative AI Profile (NIST AI 600-1)

    A companion to the AI Risk Management Framework, with actions for generative AI.

    Source 5National Institute of Standards and Technology

Sources. Where the facts come from.

Numbered as they are cited in the text. Each link opens the original.

  1. LLM06:2025 Excessive Agency

    OWASP Gen AI Security Project · 2025

    Back to the text

  2. LLM01:2025 Prompt Injection

    OWASP Gen AI Security Project · 2025

    Back to the text

  3. Guidelines for Human-AI Interaction

    Amershi et al., Microsoft Research (CHI 2019) · May 2019

    Back to the text

Let’s build what happens next.

Tell us what you’re building. We’ll answer straight.

Book a discovery call

Three ways to start

  1. 01About 2 minutes

    A quick question

    You get A reply from a lead, not a sales queue

  2. 02About 8 minutesMost useful

    A project brief

    You get Options and a first scope after one call

  3. 03About 15 minutes

    A formal RFQ or RFP

    You get Receipt confirmed and a named bid lead

Every engagement starts with a written scope and a quote agreed before work begins. How each package is priced