← The journal
Strategy

Approval Workflows for AI Agents That Auditors Can Work With

How to design AI agent approval workflows an auditor can work with, and what Swfte builds today.

Swfte Journal / Strategy

An approval workflow that an auditor can work with answers six questions without anyone having to reconstruct events from memory: who approved, were they allowed to, what did they see, when, what happened next, and can you show the record has not been altered. That is the whole test. A button labelled "Approve" that lets an agent carry on proves none of it. This post sets out how to design gates that answer those questions, which parts of the Swfte platform build them today, and which parts are still design intent and so depend on how you assign work and run your process. For the wider picture see AI governance and the platform's governance layer.

What does an auditor actually ask?

Auditors sample. They pick a handful of approvals from a period and pull on each one. The questions are plain, and they repeat.

  • Who approved? A named person, not a team mailbox and not "the system".
  • Were they allowed to? Was this person entitled to decide this class of action at that time?
  • What did they see? The exact action, the data it touched and the risk attached, at the moment of the decision.
  • When? A timestamp that is part of the record, not added later.
  • What happened next? Did the agent proceed, stop or take another branch, and is that visible in the same trail?
  • Is the record intact? Can you show that events were not edited or removed afterwards?

Notice that most of these are about the record, not the decision. A good decision with no evidence is a bad audit result. A modest decision with a complete trail is easy to defend. Design for the trail first.

What should an approval gate look like?

Start from the principle that an approval is a durable object with an owner and a lifecycle, not a pop-up. A well-formed gate has the following properties.

It is addressed to a named assignee. Someone specific is asked, so someone specific is accountable. A gate sent to "everyone" is answered by no one.

It nudges, then escalates, then parks. People are on leave, in meetings, asleep. If the assignee does not answer, the gate reminds them, then escalates to a named fallback, then parks. Parked means waiting safely, not proceeding.

It never auto-approves. This is the property auditors care about most. Silence must never be read as consent. If a gate times out, the safe outcome is a denial or a parked run, never an approval.

The first decision wins, and the decider is the authenticated user. Two people answering the same request must not both succeed, and the name on the decision must come from the login session, not from a field someone typed.

Every transition is written down. Asked, nudged, escalated, decided, expired. Each is an event with a time, so the story of the approval can be read back in order.

Which of these does the platform build today?

Several parts of the platform handle approvals, and they are not yet one unified model. Be clear about that when you plan: a customer should not assume a single inbox that covers everything. The behaviours below are built, each in the part of the platform named.

  • Studio relay gates. A gate is addressed to an assignee, nudged, escalated to a named fallback and parked. The design rule is that a gate never auto-approves. Every state change writes an immutable audit event. Gates sit alongside the workflow human input step, which pauses a run, carries an assignee and a timeout, offers approve and reject branches, and consumes each decision once, so one answer cannot quietly approve a later gate.
  • Nexus approvals for coding agents. When a coding agent wants to run a tool, a card shows the tool, a redacted target, a risk level, the repository and the branch. The first decision wins and the second gets a conflict response. The decider is stamped from the authenticated principal, and an unanswered request expires. See Nexus.
  • Policy engine asks. When a policy returns an ask verdict, the request becomes a durable pause with a bounded wait, and the rule's timeout behaviour defaults to deny.
  • Cortex desktop approvals. In Cortex, when a Claude Code session wants to do anything beyond reading, it asks, the prompt times out into a deny, and the approval is bound to the exact call that was shown. Send, pay, delete, push and shell actions always need a person. More on this on the approvals and governance page.

A caution about scope. Policy enforcement is live only for runs that have a policy attached, so part of your checklist is confirming that your agent's runs have one. And one thing we do not describe as an approval at all: Workers are headless and advisory, and their approval mode setting strips tools at deploy time rather than asking a person. Do not count it as a human gate.

Can you show the record has not been altered?

This is where a hash chain earns its place. The platform keeps a run ledger: an append-only list of events for each run, covering policy decisions, approvals, gates and more, where each event carries a hash that includes the one before it. If someone edits or removes an event in the middle, the chain stops verifying. A read endpoint returns the events for a run and a verify endpoint checks the chain.

Two honest limits. First, a chain on its own detects tampering with individual events but cannot show that a whole log was replaced wholesale. The platform has an opt-in feature that seals the head of the chain in write-once storage to close that gap, and you should ask whether it is switched on in your deployment rather than assume. Second, the ledger is readable as JSON through the API; there is not yet a dedicated export for handing to an auditor, so plan on extracting and presenting it yourself. For the broader observability picture see AI observability for agents and workflows.

What is not built?

Three things an auditor may reasonably expect are not enforced by the platform today. We would rather say so here than let you find out in a test.

  1. Approver differs from requester. There is no check that stops the person who triggered a run from also approving it. Segregation of duties is design intent. Today you get it by addressing the gate to a different named person and by reviewing the ledger to confirm that the decider differs.
  2. Approver roles and delegation. Approvals go to a named assignee. A model of roles, who may approve what, and a delegation chain for leave cover is not built. The "were they allowed to?" question is answered by your assignment and your process.
  3. A records retention policy. There is no policy object that says how long approval records are kept. The retention period is yours to decide with your records team. A library for erasing ledger payloads while keeping the chain exists, but it is not exposed as a feature you can switch on.

None of these makes approvals worthless. It means that for these three, the control lives in how you assign gates and run the process, and the platform provides the evidence to check that you did.

How does this connect to the regulations?

Treat frameworks as lists of things you may be asked to show. Approval records and gate evidence support evidence for several of them. The EU AI Act's record-keeping and human oversight requirements for high-risk systems are the obvious pair; GDPR accountability asks you to demonstrate how decisions about personal data were controlled; NIS2 and DORA both lean on governance and incident evidence; and control families in ISO 27001 and SOC 2 around access and change management are the sort of areas where auditors sample approvals. The platform supports evidence for what a customer must show. It does not decide whether a system meets a legal requirement, and no approval feature does that alone. The exact posture depends on your use case, jurisdiction, deployment and configuration, and none of this is legal advice. For a worked structure see building an AI Act evidence pack and the EU AI Act page.

Does an automated reviewer count as approval?

No, and it is worth being blunt about it because the confusion is common. An LLM reviewer that checks a draft is automated quality assurance. It is useful, but it is not a person taking responsibility. At Swfte, our own content pipelines built as Studio workflows use reviewer agents as their gates, and those pipelines have no human step. We call that QA, not approval, and we would not present it to an auditor as oversight. A worked example of a Swfte approval routed through the platform is not published yet.

What we do run on ourselves is more modest and process-based: steps that a named owner decides, with the decision written down. Examples are a paid test charge kept by the owner's explicit choice and a signing key that only the founder holds. Those are practices, not product features.

What is a sensible design checklist?

  • Name the assignee and the fallback for every gate before you build it.
  • Decide what the approver must see, and confirm the gate shows it.
  • Choose the timeout outcome in advance: deny or park, never approve.
  • Wire a deny branch, or accept that a denial cancels the run.
  • Attach a policy to the run and confirm enforcement is on.
  • Test the ledger: run the verify call and keep the output with the sample.
  • Write down how you enforce approver separation, because the platform does not yet.
  • Agree retention with your records team and record the decision.

Teams moving from pilots into this kind of operation may also like from AI pilot to governed AI operations, and the governed agents page and compliance workflows page show where this is heading.

What is built, and what is design intent?

Built in the product: addressed, nudged, escalated and parked gates; a human input step in workflows; first-decision-wins approvals with the decider taken from the authenticated user; timeouts that deny rather than approve; and a hash-chained run ledger with a verify call. Designed for: approver separation enforced by the platform, approver roles and delegation, a retention policy object, and one approval model across all products. In use at Swfte: owner-decided steps and a habit of writing the decision down. Cortex does not front our own website work today.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

See what your agents are actually doing

Nexus gives you governance, observability and spend control across every agent you run.