Short answer
Put a person in front of actions that are hard to undo, move money or data outside your systems, or change permissions, and nothing else. Pause the agent at that point, show the reviewer the intent, arguments and source, deny if nobody answers in time, and log every decision. LangGraph provides interrupt() and a human-in-the-loop middleware for the pause and resume.
The steps at a glance
- Sort every tool into approve-first or run-freely
- Set thresholds and a default for silence
- Pause the agent at the gate and resume on the decision
- Or gate at the tool level with the middleware
- Show the reviewer what they need to decide
- Record every decision, including timeouts
- Measure approvals and fix rubber-stamping
Before you start
Who this is for
- Developers adding approval steps to agents that call tools with side effects.
- Product and risk owners who must decide what an agent can do alone.
- Teams whose approvals are being clicked through without reading.
Probably not for you if
- Agents that only read and summarise with no write tools. Approval adds friction for little gain.
- Teams without an agent framework yet: start with how to build an AI agent.
Prerequisites
- An agent with tools, built in LangGraph or LangChain for the code examples (the design steps apply to any framework).
- A list of the agent's tools split into read-only and write-capable.
- Named reviewers, and a place they already work, such as a ticket queue or chat tool.
- A checkpointer for the pause and resume example. The LangGraph docs use the in-memory one for development; use a persistent one in production.
- Time
- About half a day for the design, one to two days to wire in and test
- Cost
- Free libraries. Reviewer time is the real cost: measure it in step 7.
- Skill
- Python, plus familiarity with your agent framework
Estimates are ours, not measurements, and move with your hardware, data and network.
Synchronous or asynchronous approval
Choose the pattern per action. Interactive agents where the user is present can ask in the conversation. Background agents need a queue, because the person who can approve may not be at their desk.
| Synchronous (in the session) | Asynchronous (queue) | |
|---|---|---|
| Best for | Chat assistants, coding agents with the user watching | Scheduled or event-driven agents, finance and legal actions |
| Who approves | The user in the session | A named role, possibly someone other than the requester |
| State | Held for minutes | Must persist for hours or days (persistent checkpointer) |
| Timeout default | Deny after a short wait | Deny after the queue deadline, notify the requester |
| Main risk | Reflex clicking | Requests nobody sees |
Step 1Sort every tool into approve-first or run-freely
You end up with: A table of tools with a decision for each, signed off by the agent's owner.
Start from the tool list, not from a feeling. For each tool ask four questions. Can the effect be undone? Does it move data or money outside the system? Does it speak to a customer or the public? Does it change who can do what? A yes to any of them puts the tool on the approve-first list.
Everything else runs freely, and stays logged. Keep the approve-first list short on purpose. OWASP's guidance on excessive agency says to use human-in-the-loop control to require a human to approve high-impact actions, which is a narrower instruction than "approve everything". If every call asks, people stop reading the question.
Record the decision with the owner. The table below is a sensible starting set. Adjust it to your domain, and write the threshold values with the people who own the risk.
A starting classification Action type Decision Why Read internal data the user may already see Run freely, log Low impact; approval adds noise Draft a message or document Run freely Nothing leaves until a person sends it Send an external message Approve first Cannot be recalled Write or delete in a system of record Approve first Hard to undo Spend or commit money above a threshold Approve first Financial impact; set the threshold with finance Change access or permissions Approve first Widens what the agent or others can do Export data outside the organisation Approve first Exfiltration risk Checked against: OWASP GenAI: LLM06:2025 Excessive Agency
Step 2Set thresholds and a default for silence
You end up with: Written rules for when approval is needed, who may give it, and what happens if nobody does.
Some actions need approval only above a limit: a refund over an amount, a purchase over a value, an export over a number of records. Put the limit in configuration, with an owner, and review it on a schedule. Do not bury it in the prompt, where the model could talk itself past it.
Say who may approve. The person who triggered the run is often the wrong approver for high-impact actions, because they have the same incentive as the agent to finish. For the riskiest actions require someone other than the requester, or two people.
Decide what happens on silence. The safe default is to deny after a timeout, record the denial and tell the requester. An action that proceeds because nobody answered is not an approval. Set timeouts that fit the work: minutes for an interactive session, hours for a queue.
Approval policy as configuration (our template) · yaml approvals: send_external_message: required: true approvers: [team-lead] timeout_minutes: 120 on_timeout: deny issue_refund: required_above: 100 # set with finance approvers: [finance-reviewer] different_from_requester: true timeout_minutes: 240 on_timeout: deny read_customer_record: required: false log: trueChecked against: OWASP GenAI: LLM06:2025 Excessive Agency
Step 3Pause the agent at the gate and resume on the decision
You end up with: The agent stops before the risky action, keeps its state, and continues only after a decision arrives.
LangGraph's interrupt mechanism is built for this. You call
interrupt()inside a node where you want to pause. The documentation lists three requirements: a checkpointer to persist graph state, a thread ID in the config, and the interrupt call at the point to pause. You resume by invoking the graph again withCommand(resume=...)on the same thread.There are rules that matter for approvals. When execution resumes, the runtime restarts the whole node from the beginning, not from the line where
interruptwas called, so any side effect that comes beforeinterrupt()in the node must be idempotent. Do not wrapinterrupt()in a baretry/except, do not skip it conditionally in ways that change between runs, and do not reorder several interrupts in one node.The first example below is the approve-or-cancel pattern from the documentation. Put the risky tool call in the node that the "proceed" branch leads to, so it cannot run before the decision.
Pause for a decision and route on it (from the LangGraph documentation) · python from typing import Literal from langgraph.types import Command, interrupt def approval_node(state: State) -> Command[Literal["proceed", "cancel"]]: is_approved = interrupt({ "question": "Do you want to proceed?", "details": state["action_details"] }) if is_approved: return Command(goto="proceed") else: return Command(goto="cancel")Compile with a checkpointer (development only; use a persistent one in production) · python from langgraph.checkpoint.memory import InMemorySaver checkpointer = InMemorySaver() graph = builder.compile(checkpointer=checkpointer)Resume on the same thread with the human's answer · python from langgraph.types import Command config = {"configurable": {"thread_id": "thread-1"}} resumed = graph.stream_events(Command(resume=True), config=config, version="v3")What the documentation says to expect
The first run stops at the interrupt and reports it (stream.interrupted is true, stream.interrupts holds the payload). The resumed run continues from that node.Checked against: LangGraph documentation: Interrupts
Step 4Or gate at the tool level with the middleware
You end up with: Specific tools pause for approve, edit or reject decisions without custom graph nodes.
If you build with LangChain agents,
HumanInTheLoopMiddlewarelets you declare which tools need a decision. The documentation's example setswrite_fileto allow all decisions,execute_sqlto allow only approve and reject, andread_datato need no approval, with a checkpointer on the agent.The decision types are: approve (run the tool with the arguments the agent proposed), edit (change the arguments before it runs), reject (skip the call and return rejection feedback to the agent) and respond (return the human's message as a synthetic tool result, for ask-the-user tools). Restrict the list per tool. For a SQL tool, allowing "edit" lets a reviewer rewrite a query, which may be what you want or may not.
Resume with the decisions list in a
Command. The model name in the documentation's example is just its sample; use your own.Per-tool approval configuration (from the LangChain documentation; model id is the docs' sample) · python from langchain.agents.middleware import HumanInTheLoopMiddleware from langgraph.checkpoint.memory import InMemorySaver agent = create_agent( model="gpt-5.5", tools=[write_file, execute_sql, read_data], middleware=[ HumanInTheLoopMiddleware( interrupt_on={ "write_file": True, # All decisions allowed "execute_sql": {"allowed_decisions": ["approve", "reject"]}, "read_data": False, # No approval needed }, description_prefix="Tool execution pending approval", ), ], checkpointer=InMemorySaver(), )Resume with an approve decision · python from langgraph.types import Command agent.invoke( Command( resume={"decisions": [{"type": "approve"}]} ), config=config, version="v2", )Checked against: LangChain documentation: Human-in-the-loop
Step 5Show the reviewer what they need to decide
You end up with: Each request carries the intent, arguments, evidence and consequences on one screen.
A bare "Allow this tool?" produces a reflex click. Give the reviewer: what the agent is trying to do in plain words, the exact tool and arguments, the user request that started the run, the text or data that led the agent here, what happens if approved, and what happens if denied. If the agent read untrusted content before asking, say so and show it. That is exactly the situation a prompt injection creates, as described in how to stop prompt injection.
Make the safe choice easy and the unsafe one deliberate. Reject should be one click and should let the reviewer write a reason, which goes back to the agent as feedback. Edit should show a diff of arguments. Avoid pre-selecting approve.
Put requests where reviewers already work. A separate dashboard that nobody opens means every request times out and is denied.
Fields on an approval request Field Why Agent and requester identity Who is asking, for whom Action in plain words So a non-specialist can judge it Tool name and exact arguments What will actually run Triggering request and source text Shows whether the instruction came from the user or from content Risk reason and threshold hit Why this needed approval Consequence of approve and of deny Makes the trade-off visible Expiry time Shows the deadline and the default on timeout Step 6Record every decision, including timeouts
You end up with: A durable record that shows who approved what, when, on which evidence.
Write one record per request, whatever the outcome: approved, edited, rejected, timed out. Include the request ID, agent identity, tool and arguments (redacted where they hold personal data), the evidence shown, the reviewer, the decision, any edited arguments, any reason given, timestamps for request and decision, and the policy version that required it.
Store it separately from application logs, with write-once or append-only controls where you can. These records are what an auditor or an incident reviewer will ask for, and they let you measure the process in the next step. For oversight of high-risk systems under EU rules, see the EU AI Act high-risk checklist for how this evidence is used.
An approval record (our example) · json { "request_id": "apr_0001", "agent": "procurement-agent-001", "tool": "create_purchase_order", "arguments_redacted": {"supplier_id": "S-1042", "amount": 4800}, "policy_version": "approvals-v3", "requested_at": "2026-10-06T09:14:03Z", "reviewer": "finance-reviewer-7", "decision": "approve", "reason": "", "decided_at": "2026-10-06T09:21:40Z" }Checked against: EU AI Act, Article 14: Human oversight (artificialintelligenceact.eu)
Step 7Measure approvals and fix rubber-stamping
You end up with: Numbers that show whether people are really reviewing, and a process for adjusting the gate.
Track the share of requests approved without edit, the median time to decide, the share that time out, and how often reviewers deny or edit. A near-100 per cent approval rate with decisions in a couple of seconds suggests the gate is noise or reviewers are not looking. A high timeout rate means the queue is in the wrong place or the wrong people are assigned.
Seed some known-bad requests into the stream, clearly recorded as tests, and see whether reviewers catch them. Treat misses as a design problem: shorten the queue, improve the display, or remove approval from low-risk actions to bring attention back to the rest.
Review the gate monthly. If an action is approved every time for several months with no edits and little impact, consider whether to move it to run-freely with monitoring. If an incident shows an ungated action caused harm, add it to the approve-first list and record why.
When approval is the wrong control
Approval is a weak fix for an agent that has too much access. If the agent can read every customer record and send email, a person reviewing each call is doing work a permission could do. Cut the tools and scopes first, as described in how to govern AI agents, and keep human approval for the actions that remain high-impact.
It is also not a defence against being deceived. A reviewer shown a convincing request built from injected text may approve it. Show the source and keep the list of approve-first actions short enough that people stay alert.
Troubleshooting
| What you see | Likely cause | Fix |
|---|---|---|
| The run raises an error about a missing checkpointer or thread ID when it hits interrupt() | interrupt() needs a checkpointer and a thread ID in the config. | Compile the graph with a checkpointer and pass {"configurable": {"thread_id": "..."}} on every call, using the same ID to resume. |
| A side effect, such as sending a message, happens twice after resume | The runtime restarts the whole node on resume, so code before interrupt() runs again. | Move side effects after the interrupt or make them idempotent, with an idempotency key. |
| Resuming does nothing, or starts a fresh run | The resume call used a different thread ID, or the in-memory checkpointer was lost on restart. | Reuse the original thread ID and use a persistent checkpointer for anything that outlives one process. |
| Reviewers approve almost everything within seconds | Too many requests, low-risk actions in the queue, or a display with no context. | Shrink the approve-first list, add evidence fields and run a test with seeded bad requests. |
| Requests expire unseen and agents stall | The queue is in a tool reviewers do not open, or no one is assigned out of hours. | Send requests to the reviewers' existing channel, name a backup approver, and alert on queue age. |
| The agent retries a rejected action in a different form | Rejection feedback does not tell the model not to try again, and nothing blocks the alternative. | Return a clear rejection reason, block the specific tool for that run after a rejection, and log repeated attempts. |
Verify it worked
Next steps
- How to govern AI agents: where approval sits alongside identity, policy and autonomy levels
- How to stop prompt injection: why reviewers need to see the source of an instruction
- Controlled autonomy: a way to decide how much an agent does alone
- What to automate first with approval gates: a worked example in security operations
Related guides
- How to Govern AI Agents: Identity, Policy, Approvals: Govern agents at runtime: list every agent, give each an identity and an owner, write down what it may and may not do in a Trust Profile, enforce allow, deny and approve rules, choose an autonomy level, record every action and review on a schedule.
- How to Stop Prompt Injection: Layered Defences That Work: A layered defence for prompt injection: assume it will happen, keep untrusted content apart from instructions, give agents the fewest tools and shortest-lived privileges, require human approval for risky actions, block data leaving, and test with canary documents.
- How to Secure MCP Servers: OAuth, Scopes, Allow-Lists: Harden an MCP deployment against the attacks the specification names: validate token audience, never pass tokens through, ask for minimal scopes, sandbox local servers, allow-list servers, gate sensitive tools with a human, and log every call.
- How to Monitor AI Agents in Production (2026 Guide): Give every agent run an id, record each model and tool step as a span, redact before you store, alert on loops, tool failures and cost per run, and read a weekly sample by hand.
- How to Build an AI Agent in 2026 (Code & No-Code): Build an agent as a loop with tools, a turn budget and an eval set: define a narrow job, pick a build path, connect tools, choose models per step, test, deploy and observe.
Frequently asked questions
When should an AI agent ask for human approval?
When an action is hard to undo, moves money or data outside your systems, contacts people outside the organisation, or changes permissions. Read-only and draft-only actions can usually run freely with logging. Keep the approve-first list short so reviewers keep paying attention.
How do I add human-in-the-loop to a LangGraph agent?
Call interrupt() in the node where the agent should pause, compile the graph with a checkpointer, run with a thread ID, and resume with Command(resume=...) on the same thread. Remember that the node restarts from the beginning on resume, so earlier side effects must be idempotent.
What should happen if nobody responds to an approval request?
Deny it, record the denial and tell the requester. An action that proceeds because nobody answered is not approved. Set the timeout to match the work: minutes for an interactive session, hours for a queue, and name a backup approver.
What is approval fatigue and how do I prevent it?
It is reviewers clicking approve without reading because they see too many requests with too little context. Reduce requests to the high-impact actions, show intent, arguments and source, make reject easy, and measure approval time and rate. Seed known-bad tests to see whether people catch them.
Is human approval required by the EU AI Act?
For high-risk AI systems, Article 14 says they must be designed so natural persons can effectively oversee them while in use, including being able to override the output or stop the system, and to stay aware of automation bias. It does not prescribe approving every action. This is not legal advice: check the text for your case.
Should the person who started the agent also approve its actions?
Not for high-impact actions. The requester has the same goal as the agent. Require a different reviewer for the riskiest actions, or two reviewers, and record who approved. For low-risk actions in an interactive session, the user approving their own request is reasonable.
How Swfte can help
Swfte describes approvals as one of six governance decisions and a step on its autonomy scale. The pages below explain that model and an example from security operations.
- Controlled autonomy: five levels from assist to adaptive, with approval at level two
- AI SOC agents: a worked case of approval gates in security operations
- Trust and governance fabric: where approval fits in the wider governance model
How approvals are configured in a Swfte deployment depends on the product and plan: <approval workflow availability - founder to fill>. You can build this pattern with LangGraph or another framework without Swfte.
Missing a step or found a command that no longer works? Tell us, or request a how-to.
Sources and last verified
Commands, versions and facts in this guide were checked against the sources below on . Tools change quickly: if something differs from what you see, trust the official documentation and let us know.
- EU AI Act, Article 14: Human oversight (artificialintelligenceact.eu): Paragraph 1 oversight requirement for high-risk systems; paragraph 4 items on awareness of automation bias, overriding output and a stop procedure.
- LangGraph documentation: Interrupts: interrupt() and Command(resume=...) code, checkpointer and thread ID requirements, restart-from-node-start behaviour, rules on try/except and idempotent side effects.
- LangChain documentation: Human-in-the-loop: HumanInTheLoopMiddleware import path, interrupt_on configuration, approve, edit, reject and respond decision types, resume command.
- OWASP GenAI: LLM06:2025 Excessive Agency: Human-in-the-loop for high-impact actions, minimising extensions and permissions, authorisation in downstream systems.
- Microsoft Learn: Defend against indirect prompt injection attacks: Human-in-the-loop described as the last line of defence for risky actions; page dated 2026-03-19.
Topics
- human in the loop
- approvals
- agents
- LangGraph
- oversight
Machine-readable copies: this guide as markdown, index of all guides (JSON). Canonical address: https://www.swfte.com/how-to-set-up-human-approval-for-ai-agents.