Govern · Intermediate

How to set up human approval for AI agents

  • Time: About half a day for the design, one to two days to wire in and test
  • Cost: Free libraries. Reviewer time is the real cost: measure it in step 7.
  • Level: Intermediate
On this page
  1. Short answer
  2. Before you start
  3. Synchronous or asynchronous approval
  4. 1. Sort every tool into approve-first or run-freely
  5. 2. Set thresholds and a default for silence
  6. 3. Pause the agent at the gate and resume on the decision
  7. 4. Or gate at the tool level with the middleware
  8. 5. Show the reviewer what they need to decide
  9. 6. Record every decision, including timeouts
  10. 7. Measure approvals and fix rubber-stamping
  11. When approval is the wrong control
  12. Troubleshooting
  13. Verify it worked
  14. Next steps
  15. FAQ
  16. How Swfte can help
  17. Sources and last verified

Short answer

Put a person in front of actions that are hard to undo, move money or data outside your systems, or change permissions, and nothing else. Pause the agent at that point, show the reviewer the intent, arguments and source, deny if nobody answers in time, and log every decision. LangGraph provides interrupt() and a human-in-the-loop middleware for the pause and resume.

The steps at a glance

  1. Sort every tool into approve-first or run-freely
  2. Set thresholds and a default for silence
  3. Pause the agent at the gate and resume on the decision
  4. Or gate at the tool level with the middleware
  5. Show the reviewer what they need to decide
  6. Record every decision, including timeouts
  7. Measure approvals and fix rubber-stamping

Before you start

Who this is for

  • Developers adding approval steps to agents that call tools with side effects.
  • Product and risk owners who must decide what an agent can do alone.
  • Teams whose approvals are being clicked through without reading.

Probably not for you if

  • Agents that only read and summarise with no write tools. Approval adds friction for little gain.
  • Teams without an agent framework yet: start with how to build an AI agent.

Prerequisites

  • An agent with tools, built in LangGraph or LangChain for the code examples (the design steps apply to any framework).
  • A list of the agent's tools split into read-only and write-capable.
  • Named reviewers, and a place they already work, such as a ticket queue or chat tool.
  • A checkpointer for the pause and resume example. The LangGraph docs use the in-memory one for development; use a persistent one in production.
Time
About half a day for the design, one to two days to wire in and test
Cost
Free libraries. Reviewer time is the real cost: measure it in step 7.
Skill
Python, plus familiarity with your agent framework

Estimates are ours, not measurements, and move with your hardware, data and network.

Synchronous or asynchronous approval

Choose the pattern per action. Interactive agents where the user is present can ask in the conversation. Background agents need a queue, because the person who can approve may not be at their desk.

Synchronous (in the session)Asynchronous (queue)
Best forChat assistants, coding agents with the user watchingScheduled or event-driven agents, finance and legal actions
Who approvesThe user in the sessionA named role, possibly someone other than the requester
StateHeld for minutesMust persist for hours or days (persistent checkpointer)
Timeout defaultDeny after a short waitDeny after the queue deadline, notify the requester
Main riskReflex clickingRequests nobody sees
  1. Step 1Sort every tool into approve-first or run-freely

    You end up with: A table of tools with a decision for each, signed off by the agent's owner.

    Start from the tool list, not from a feeling. For each tool ask four questions. Can the effect be undone? Does it move data or money outside the system? Does it speak to a customer or the public? Does it change who can do what? A yes to any of them puts the tool on the approve-first list.

    Everything else runs freely, and stays logged. Keep the approve-first list short on purpose. OWASP's guidance on excessive agency says to use human-in-the-loop control to require a human to approve high-impact actions, which is a narrower instruction than "approve everything". If every call asks, people stop reading the question.

    Record the decision with the owner. The table below is a sensible starting set. Adjust it to your domain, and write the threshold values with the people who own the risk.

    A starting classification
    Action typeDecisionWhy
    Read internal data the user may already seeRun freely, logLow impact; approval adds noise
    Draft a message or documentRun freelyNothing leaves until a person sends it
    Send an external messageApprove firstCannot be recalled
    Write or delete in a system of recordApprove firstHard to undo
    Spend or commit money above a thresholdApprove firstFinancial impact; set the threshold with finance
    Change access or permissionsApprove firstWidens what the agent or others can do
    Export data outside the organisationApprove firstExfiltration risk

    Checked against: OWASP GenAI: LLM06:2025 Excessive Agency

  2. Step 2Set thresholds and a default for silence

    You end up with: Written rules for when approval is needed, who may give it, and what happens if nobody does.

    Some actions need approval only above a limit: a refund over an amount, a purchase over a value, an export over a number of records. Put the limit in configuration, with an owner, and review it on a schedule. Do not bury it in the prompt, where the model could talk itself past it.

    Say who may approve. The person who triggered the run is often the wrong approver for high-impact actions, because they have the same incentive as the agent to finish. For the riskiest actions require someone other than the requester, or two people.

    Decide what happens on silence. The safe default is to deny after a timeout, record the denial and tell the requester. An action that proceeds because nobody answered is not an approval. Set timeouts that fit the work: minutes for an interactive session, hours for a queue.

    Approval policy as configuration (our template) · yaml
    approvals:
      send_external_message:
        required: true
        approvers: [team-lead]
        timeout_minutes: 120
        on_timeout: deny
      issue_refund:
        required_above: 100      # set with finance
        approvers: [finance-reviewer]
        different_from_requester: true
        timeout_minutes: 240
        on_timeout: deny
      read_customer_record:
        required: false
        log: true

    Checked against: OWASP GenAI: LLM06:2025 Excessive Agency

  3. Step 3Pause the agent at the gate and resume on the decision

    You end up with: The agent stops before the risky action, keeps its state, and continues only after a decision arrives.

    LangGraph's interrupt mechanism is built for this. You call interrupt() inside a node where you want to pause. The documentation lists three requirements: a checkpointer to persist graph state, a thread ID in the config, and the interrupt call at the point to pause. You resume by invoking the graph again with Command(resume=...) on the same thread.

    There are rules that matter for approvals. When execution resumes, the runtime restarts the whole node from the beginning, not from the line where interrupt was called, so any side effect that comes before interrupt() in the node must be idempotent. Do not wrap interrupt() in a bare try/except, do not skip it conditionally in ways that change between runs, and do not reorder several interrupts in one node.

    The first example below is the approve-or-cancel pattern from the documentation. Put the risky tool call in the node that the "proceed" branch leads to, so it cannot run before the decision.

    Pause for a decision and route on it (from the LangGraph documentation) · python
    from typing import Literal
    from langgraph.types import Command, interrupt
    
    def approval_node(state: State) -> Command[Literal["proceed", "cancel"]]:
        is_approved = interrupt({
            "question": "Do you want to proceed?",
            "details": state["action_details"]
        })
        if is_approved:
            return Command(goto="proceed")
        else:
            return Command(goto="cancel")
    Compile with a checkpointer (development only; use a persistent one in production) · python
    from langgraph.checkpoint.memory import InMemorySaver
    
    checkpointer = InMemorySaver()
    graph = builder.compile(checkpointer=checkpointer)
    Resume on the same thread with the human's answer · python
    from langgraph.types import Command
    
    config = {"configurable": {"thread_id": "thread-1"}}
    resumed = graph.stream_events(Command(resume=True), config=config, version="v3")

    What the documentation says to expect

    The first run stops at the interrupt and reports it (stream.interrupted is true, stream.interrupts holds the payload). The resumed run continues from that node.

    Checked against: LangGraph documentation: Interrupts

  4. Step 4Or gate at the tool level with the middleware

    You end up with: Specific tools pause for approve, edit or reject decisions without custom graph nodes.

    If you build with LangChain agents, HumanInTheLoopMiddleware lets you declare which tools need a decision. The documentation's example sets write_file to allow all decisions, execute_sql to allow only approve and reject, and read_data to need no approval, with a checkpointer on the agent.

    The decision types are: approve (run the tool with the arguments the agent proposed), edit (change the arguments before it runs), reject (skip the call and return rejection feedback to the agent) and respond (return the human's message as a synthetic tool result, for ask-the-user tools). Restrict the list per tool. For a SQL tool, allowing "edit" lets a reviewer rewrite a query, which may be what you want or may not.

    Resume with the decisions list in a Command. The model name in the documentation's example is just its sample; use your own.

    Per-tool approval configuration (from the LangChain documentation; model id is the docs' sample) · python
    from langchain.agents.middleware import HumanInTheLoopMiddleware
    from langgraph.checkpoint.memory import InMemorySaver
    
    agent = create_agent(
        model="gpt-5.5",
        tools=[write_file, execute_sql, read_data],
        middleware=[
            HumanInTheLoopMiddleware(
                interrupt_on={
                    "write_file": True,  # All decisions allowed
                    "execute_sql": {"allowed_decisions": ["approve", "reject"]},
                    "read_data": False,  # No approval needed
                },
                description_prefix="Tool execution pending approval",
            ),
        ],
        checkpointer=InMemorySaver(),
    )
    Resume with an approve decision · python
    from langgraph.types import Command
    
    agent.invoke(
        Command(
            resume={"decisions": [{"type": "approve"}]}
        ),
        config=config,
        version="v2",
    )

    Checked against: LangChain documentation: Human-in-the-loop

  5. Step 5Show the reviewer what they need to decide

    You end up with: Each request carries the intent, arguments, evidence and consequences on one screen.

    A bare "Allow this tool?" produces a reflex click. Give the reviewer: what the agent is trying to do in plain words, the exact tool and arguments, the user request that started the run, the text or data that led the agent here, what happens if approved, and what happens if denied. If the agent read untrusted content before asking, say so and show it. That is exactly the situation a prompt injection creates, as described in how to stop prompt injection.

    Make the safe choice easy and the unsafe one deliberate. Reject should be one click and should let the reviewer write a reason, which goes back to the agent as feedback. Edit should show a diff of arguments. Avoid pre-selecting approve.

    Put requests where reviewers already work. A separate dashboard that nobody opens means every request times out and is denied.

    Fields on an approval request
    FieldWhy
    Agent and requester identityWho is asking, for whom
    Action in plain wordsSo a non-specialist can judge it
    Tool name and exact argumentsWhat will actually run
    Triggering request and source textShows whether the instruction came from the user or from content
    Risk reason and threshold hitWhy this needed approval
    Consequence of approve and of denyMakes the trade-off visible
    Expiry timeShows the deadline and the default on timeout
  6. Step 6Record every decision, including timeouts

    You end up with: A durable record that shows who approved what, when, on which evidence.

    Write one record per request, whatever the outcome: approved, edited, rejected, timed out. Include the request ID, agent identity, tool and arguments (redacted where they hold personal data), the evidence shown, the reviewer, the decision, any edited arguments, any reason given, timestamps for request and decision, and the policy version that required it.

    Store it separately from application logs, with write-once or append-only controls where you can. These records are what an auditor or an incident reviewer will ask for, and they let you measure the process in the next step. For oversight of high-risk systems under EU rules, see the EU AI Act high-risk checklist for how this evidence is used.

    An approval record (our example) · json
    {
      "request_id": "apr_0001",
      "agent": "procurement-agent-001",
      "tool": "create_purchase_order",
      "arguments_redacted": {"supplier_id": "S-1042", "amount": 4800},
      "policy_version": "approvals-v3",
      "requested_at": "2026-10-06T09:14:03Z",
      "reviewer": "finance-reviewer-7",
      "decision": "approve",
      "reason": "",
      "decided_at": "2026-10-06T09:21:40Z"
    }

    Checked against: EU AI Act, Article 14: Human oversight (artificialintelligenceact.eu)

  7. Step 7Measure approvals and fix rubber-stamping

    You end up with: Numbers that show whether people are really reviewing, and a process for adjusting the gate.

    Track the share of requests approved without edit, the median time to decide, the share that time out, and how often reviewers deny or edit. A near-100 per cent approval rate with decisions in a couple of seconds suggests the gate is noise or reviewers are not looking. A high timeout rate means the queue is in the wrong place or the wrong people are assigned.

    Seed some known-bad requests into the stream, clearly recorded as tests, and see whether reviewers catch them. Treat misses as a design problem: shorten the queue, improve the display, or remove approval from low-risk actions to bring attention back to the rest.

    Review the gate monthly. If an action is approved every time for several months with no edits and little impact, consider whether to move it to run-freely with monitoring. If an incident shows an ungated action caused harm, add it to the approve-first list and record why.

When approval is the wrong control

Approval is a weak fix for an agent that has too much access. If the agent can read every customer record and send email, a person reviewing each call is doing work a permission could do. Cut the tools and scopes first, as described in how to govern AI agents, and keep human approval for the actions that remain high-impact.

It is also not a defence against being deceived. A reviewer shown a convincing request built from injected text may approve it. Show the source and keep the list of approve-first actions short enough that people stay alert.

Troubleshooting

What you seeLikely causeFix
The run raises an error about a missing checkpointer or thread ID when it hits interrupt()interrupt() needs a checkpointer and a thread ID in the config.Compile the graph with a checkpointer and pass {"configurable": {"thread_id": "..."}} on every call, using the same ID to resume.
A side effect, such as sending a message, happens twice after resumeThe runtime restarts the whole node on resume, so code before interrupt() runs again.Move side effects after the interrupt or make them idempotent, with an idempotency key.
Resuming does nothing, or starts a fresh runThe resume call used a different thread ID, or the in-memory checkpointer was lost on restart.Reuse the original thread ID and use a persistent checkpointer for anything that outlives one process.
Reviewers approve almost everything within secondsToo many requests, low-risk actions in the queue, or a display with no context.Shrink the approve-first list, add evidence fields and run a test with seeded bad requests.
Requests expire unseen and agents stallThe queue is in a tool reviewers do not open, or no one is assigned out of hours.Send requests to the reviewers' existing channel, name a backup approver, and alert on queue age.
The agent retries a rejected action in a different formRejection feedback does not tell the model not to try again, and nothing blocks the alternative.Return a clear rejection reason, block the specific tool for that run after a rejection, and log repeated attempts.

Verify it worked

Next steps

Related guides

  • How to Govern AI Agents: Identity, Policy, Approvals: Govern agents at runtime: list every agent, give each an identity and an owner, write down what it may and may not do in a Trust Profile, enforce allow, deny and approve rules, choose an autonomy level, record every action and review on a schedule.
  • How to Stop Prompt Injection: Layered Defences That Work: A layered defence for prompt injection: assume it will happen, keep untrusted content apart from instructions, give agents the fewest tools and shortest-lived privileges, require human approval for risky actions, block data leaving, and test with canary documents.
  • How to Secure MCP Servers: OAuth, Scopes, Allow-Lists: Harden an MCP deployment against the attacks the specification names: validate token audience, never pass tokens through, ask for minimal scopes, sandbox local servers, allow-list servers, gate sensitive tools with a human, and log every call.
  • How to Monitor AI Agents in Production (2026 Guide): Give every agent run an id, record each model and tool step as a span, redact before you store, alert on loops, tool failures and cost per run, and read a weekly sample by hand.
  • How to Build an AI Agent in 2026 (Code & No-Code): Build an agent as a loop with tools, a turn budget and an eval set: define a narrow job, pick a build path, connect tools, choose models per step, test, deploy and observe.

Frequently asked questions

When should an AI agent ask for human approval?

When an action is hard to undo, moves money or data outside your systems, contacts people outside the organisation, or changes permissions. Read-only and draft-only actions can usually run freely with logging. Keep the approve-first list short so reviewers keep paying attention.

How do I add human-in-the-loop to a LangGraph agent?

Call interrupt() in the node where the agent should pause, compile the graph with a checkpointer, run with a thread ID, and resume with Command(resume=...) on the same thread. Remember that the node restarts from the beginning on resume, so earlier side effects must be idempotent.

What should happen if nobody responds to an approval request?

Deny it, record the denial and tell the requester. An action that proceeds because nobody answered is not approved. Set the timeout to match the work: minutes for an interactive session, hours for a queue, and name a backup approver.

What is approval fatigue and how do I prevent it?

It is reviewers clicking approve without reading because they see too many requests with too little context. Reduce requests to the high-impact actions, show intent, arguments and source, make reject easy, and measure approval time and rate. Seed known-bad tests to see whether people catch them.

Is human approval required by the EU AI Act?

For high-risk AI systems, Article 14 says they must be designed so natural persons can effectively oversee them while in use, including being able to override the output or stop the system, and to stay aware of automation bias. It does not prescribe approving every action. This is not legal advice: check the text for your case.

Should the person who started the agent also approve its actions?

Not for high-impact actions. The requester has the same goal as the agent. Require a different reviewer for the riskiest actions, or two reviewers, and record who approved. For low-risk actions in an interactive session, the user approving their own request is reasonable.

How Swfte can help

Swfte describes approvals as one of six governance decisions and a step on its autonomy scale. The pages below explain that model and an example from security operations.

How approvals are configured in a Swfte deployment depends on the product and plan: <approval workflow availability - founder to fill>. You can build this pattern with LangGraph or another framework without Swfte.

Missing a step or found a command that no longer works? Tell us, or request a how-to.

Sources and last verified

Commands, versions and facts in this guide were checked against the sources below on . Tools change quickly: if something differs from what you see, trust the official documentation and let us know.

  1. EU AI Act, Article 14: Human oversight (artificialintelligenceact.eu): Paragraph 1 oversight requirement for high-risk systems; paragraph 4 items on awareness of automation bias, overriding output and a stop procedure.
  2. LangGraph documentation: Interrupts: interrupt() and Command(resume=...) code, checkpointer and thread ID requirements, restart-from-node-start behaviour, rules on try/except and idempotent side effects.
  3. LangChain documentation: Human-in-the-loop: HumanInTheLoopMiddleware import path, interrupt_on configuration, approve, edit, reject and respond decision types, resume command.
  4. OWASP GenAI: LLM06:2025 Excessive Agency: Human-in-the-loop for high-impact actions, minimising extensions and permissions, authorisation in downstream systems.
  5. Microsoft Learn: Defend against indirect prompt injection attacks: Human-in-the-loop described as the last line of defence for risky actions; page dated 2026-03-19.

Topics

  • human in the loop
  • approvals
  • agents
  • LangGraph
  • oversight

Machine-readable copies: this guide as markdown, index of all guides (JSON). Canonical address: https://www.swfte.com/how-to-set-up-human-approval-for-ai-agents.

Build this in Studio

Describe what you need in plain language. Studio builds the agents and workflows, and you keep every version.