# How to govern AI agents

Canonical: https://www.swfte.com/how-to-govern-ai-agents
Last verified: 2026-10-06
Difficulty: Intermediate
Time: About one week for the first inventory, profiles and policy on your main agents; then a standing monthly review
Cost: No software is required to start; a spreadsheet or YAML files in a repository will do. Tools for enforcement and logging are your choice.

## Short answer

Governing an AI agent means being able to say who is acting, allowed to do what, under which policy, with which data and model, at what risk and with what oversight, and to prove it afterwards. List your agents, give each an identity and an owner, write its permissions as a profile, enforce them in the runtime, pick an autonomy level, log actions and review on a schedule.

## Who this is for

- Heads of AI, platform and security teams who have agents in pilot or production and no single list of them.
- Risk and compliance leads who must show how agent actions are controlled and evidenced.
- Engineers asked to "add governance" and unsure what that changes in the code.

Not for:
- Teams that only need content filters on a chatbot. Guardrails may be enough: see [how to stop prompt injection](https://www.swfte.com/how-to-stop-prompt-injection).
- Anyone looking for legal advice on a specific regulation. This guide is a working method, not a legal opinion.

## Prerequisites

- A rough list of where agents run: framework code, no-code builders, copilots, scheduled jobs, vendor agents inside SaaS tools.
- An identity provider or secret store where agent credentials can be issued and revoked.
- A named person on each side: a business owner and a technical owner for every agent.
- Somewhere to keep structured records: a repository, registry or governance tool.

## Guardrails and governance are different jobs

Guardrails filter content: they block a harmful output or a known attack pattern. Governance decides authority: which agent may do what, with whose data, under which policy, with whom watching. You need both. A guardrail cannot tell you whether an agent should have been allowed to send that email at all, and a policy cannot read every sentence for harm.

In practice, guardrails sit inside the model call and the tool wrapper, and governance sits in the identity, policy and records around them. If your agents have guardrails and no inventory, owners or logs, you have the first half only.

## Steps

### Step 1: Build an agent inventory

Outcome: One list of every agent, with an owner, purpose, environment and the systems it touches.

You cannot govern what you cannot list. Start with a sweep: ask each team what agents, copilots and automated assistants they run, check your identity provider and cloud accounts for service identities and API keys named after tools, and look in procurement for AI features inside software you already pay for. The discovery methods are close to those in [how to detect shadow AI](https://www.swfte.com/how-to-detect-shadow-ai).

Record each agent once. A useful minimum is: name, purpose in one sentence, business owner, technical owner, where it runs, which models it calls, which tools and systems it can reach, what data it reads and writes, and whether a person is involved before it acts. Give each agent an ID that will not change.

Do not wait for a perfect list. Publish the first version in a week, mark gaps as "unknown", and make the owner field mandatory. An agent with no owner is the first finding of the exercise.

### Step 2: Give every agent its own identity and short-lived access

Outcome: Each agent authenticates as itself, with credentials you can revoke without touching anything else.

Shared API keys and a human's personal token are the usual starting point and the worst place to end up. If five agents share one key, the log cannot say which acted, and revoking the key stops all five. Give each agent its own non-human identity, and issue credentials with a short life, tied to the task where you can.

Identity vendors now provide this as a feature. Microsoft's documentation for Entra Agent ID, for example, describes agent identities as separate from users, with the same kinds of adaptive access policy, risk detection and lifecycle management as users and workloads, and says all agent authentication and activity is logged for audit. Check the licensing on that page before you plan around it, since it names a separate licence requirement. Whatever you use, the principle is the same: one agent, one identity, one owner.

Link the identity to the inventory ID so a log line can be traced back to a named owner. See the page on [non-human identity](https://www.swfte.com/non-human-identity) for more on credentials and rotation.

### Step 3: Write a Trust Profile for each agent

Outcome: A short structured record of what the agent is, may do, may not do and what needs a human.

Turn the inventory row into a profile that a policy engine, a reviewer or an auditor can read. Swfte's Trust Profile uses these fields: identity, owner, risk level, approved models, data classification, data residency, permitted systems, allowed actions, restricted actions, human approval rule, retention, audit and policy set. You can keep the same fields in any format.

Write the "cannot" list first, because it is the part people skip. The example below is the worked example in Swfte's positioning: an AI procurement agent that can read approved supplier information, analyse contracts, compare pricing, prepare purchase recommendations and create draft purchase orders, but cannot access unrelated employee data, approve its own high-value transaction, make payments or modify restricted records. It requires approval for purchases above a threshold, contractual changes and sensitive external communications.

Set the risk level by what the agent can do, not by how clever it is. An agent that can only draft is lower risk than one that can send and pay. The risk level decides the autonomy level in step 5 and how often you review in step 8.

A Trust Profile as YAML (our template; field names follow the Swfte Trust Profile):

```yaml
identity: procurement-agent-001
owner: Procurement Operations
risk_level: medium
approved_models: [model-a, model-b]
data_classification: confidential
data_residency: EU
permitted_systems: [supplier-database, contract-store, purchasing-system-drafts]
allowed_actions:
  - read approved supplier information
  - analyse contracts
  - compare pricing
  - prepare purchase recommendations
  - create draft purchase orders
restricted_actions:
  - access unrelated employee data
  - approve its own high-value transaction
  - make payments
  - modify restricted records
human_approval:
  required_for:
    - purchases above threshold
    - contractual changes
    - sensitive external communications
retention: defined by organisational policy
audit: full action trace
policy_set: [corporate-ai-policy, supplier-data-policy]
```

> NOTE: The "threshold" is yours to set. Pick a number with finance and write it in the profile, not in a person's head.

### Step 4: Turn the profile into enforced rules

Outcome: Each requested action gets one of six decisions, enforced by the runtime rather than by a document.

A profile that nobody enforces is paperwork. Guardrails ask "should this be blocked?" Governance asks a wider set: who is acting, allowed to do what, under which policy, with which data, using which model, at what risk and with what oversight, and can we prove it. Make the decision explicit, with six outcomes: allow, deny, warn, filter, escalate, or require human approval.

Put the enforcement point where actions happen: in the tool layer or gateway that sits between the agent and your systems, not in the prompt. A prompt can be argued with. A tool call that fails an authorisation check cannot. Read the permitted systems and allowed actions from the profile, deny anything not listed, and require approval where the profile says so.

The OWASP guidance on excessive agency says the same: implement authorisation in downstream systems rather than relying on the LLM to decide whether an action is allowed, and use human-in-the-loop control for high-impact actions. Keep rules small and testable. Each rule should say who, what, on which resource, under which condition, and what the decision is. Change them through review like code, and record which policy version applied to each action.

**The six governance decisions**

| Decision | What it does | Example |
| --- | --- | --- |
| Allow | Action proceeds | Read an approved supplier record |
| Deny | Action is refused and logged | Read unrelated employee data |
| Warn | Proceeds and flags for review | Unusual volume of reads in an hour |
| Filter | Proceeds with part of the data removed | Mask personal fields in a result |
| Escalate | Passes to a named person or team | Contract clause outside the usual range |
| Require human approval | Waits for an explicit yes | Purchase above the threshold |

### Step 5: Choose an autonomy level per agent and raise it only on evidence

Outcome: Each agent runs at a stated level, with the evidence needed to move it up or down.

Autonomy is a dial, not a switch. Swfte's positioning uses five levels of controlled autonomy: L1 Assist, where the AI recommends; L2 Approve, where a human approves; L3 Supervise, where it acts within limits and is monitored; L4 Autonomous, where it acts independently within strict policy and risk bounds; and L5 Adaptive, where it improves within controlled boundaries. You can use your own scale. What matters is that each agent has a written level.

Start low. A new agent runs at L1 or L2 until it has a record. Write the promotion criteria: for example, a defined number of reviewed runs with no policy violations, an eval pass above an agreed bar, and a named owner's sign-off. Also write the demotion rule: any serious incident drops the agent a level until the cause is fixed.

Match level to risk. A read-only research agent can reasonably sit at L3. An agent that moves money should stay at L2 for the actions that spend. The detail of designing the approval step is in [how to set up human approval for AI agents](https://www.swfte.com/how-to-set-up-human-approval-for-ai-agents).

**Controlled autonomy levels**

| Level | Name | Meaning |
| --- | --- | --- |
| L1 | Assist | The AI recommends; a person acts |
| L2 | Approve | The AI prepares the action; a person approves it |
| L3 | Supervise | Acts within limits, monitored |
| L4 | Autonomous | Acts independently within strict policy and risk bounds |
| L5 | Adaptive | Improves within controlled boundaries |

### Step 6: Record who did what, under which policy, and the outcome

Outcome: A trace from data and model through decision and action to outcome, for any run.

Evidence is what turns a claim of control into proof. For every action record: agent identity, the user or trigger that started the run, data accessed, model used, output, tools called, policy applied (with version), the decision, any approval and who gave it, the action taken and the outcome. That is the record Swfte's positioning lists for the procurement example, and it is a sound general minimum.

Store logs where the agent cannot edit them, keep them for a period set by your retention policy, and make them searchable by agent ID, user and time. Redact or tokenise personal data in logs where you can, and keep access to the logs itself to a short list.

Check that the trace is complete by picking a real outcome and walking it backward: what did the agent do, why was it allowed, and who approved it? If any link is missing, fix the logging before you scale the agent.

### Step 7: Build a kill switch and test it

Outcome: You can stop one agent, or all agents, in minutes, and you have shown that it works.

Decide in advance how an agent is stopped. The mechanisms are plain: revoke its credentials, disable its entry in the tool gateway, pause its schedule, and cut its network route. Make at least two of these work without a code deployment, and make sure the on-call person knows where they are.

Separate the stop levels: pause one agent, pause a class of agents (all agents that can send email, say), pause everything. Write who may trigger each. After a stop, an agent should not restart without its owner confirming and the cause being recorded.

Run a drill. Once a quarter, pick an agent in a test environment, trigger the stop, and time it. Include the steps that people forget: queued jobs and scheduled retries that restart the agent a few minutes later.

### Step 8: Review on a schedule and map it to a framework

Outcome: A calendar of reviews and a mapping of your controls to the framework your auditors use.

Set a review cadence from the risk level: higher-risk agents more often. At each review the owner looks at the sampled runs, policy violations and denials, approvals and overrides, incidents, cost, and whether the profile still matches what the agent does. Update the profile and policy in the same session. Retire agents nobody owns.

If you need to show your approach against a framework, two are widely used. The NIST AI Risk Management Framework, which NIST describes as intended for voluntary use, is organised around four functions: Govern, Map, Measure and Manage. The ISO/IEC 42001:2023 standard specifies requirements for an artificial intelligence management system and was published in December 2023. Map your steps to them: inventory and profiles fit Map; enforcement and logs fit Manage and Measure; ownership, policy and review fit Govern.

Mapping is a convenience for people who read the standard, not a claim that you comply with it. Certification or conformity assessment is a separate process. For the audit side, see [how to audit AI systems](https://www.swfte.com/how-to-audit-ai-systems).

## Troubleshooting

| Symptom | Likely cause | Fix |
| --- | --- | --- |
| You find agents nobody claims | Teams built them as experiments, or a vendor feature switched them on. | Give owners two weeks to claim an agent in the inventory; suspend the credentials of any that stay unclaimed, and record the decision. |
| The log shows an action but not which agent did it | Agents share a credential, or a service account stands in for several. | Issue one identity per agent and log the identity from the verified credential, not from a field the agent sets. |
| Approval requests pile up and people approve without reading | Too many actions marked as needing approval, or the request lacks context. | Limit approval to the actions in the profile's approval rule, show the reviewer the intent and arguments, and measure approval time and denial rate. |
| The profile says one thing and the agent does another | Permissions in the tool layer were never tied to the profile, or have drifted since. | Generate the allowed tool list from the profile, and add a test that fails if a tool outside the profile is callable. |
| Auditors ask for evidence you cannot produce | Logs record the model output but not the policy version, approval or data accessed. | Extend the record fields in step 6 and replay a recent run end to end to confirm the trace is complete. |
| An agent's autonomy level crept up without anyone deciding | Approvals were relaxed to ease friction and nobody recorded it. | Make level changes a reviewed change with the evidence attached, and report current levels at each governance review. |

## Verify it worked

- [ ] Every agent in use appears in the inventory with a business owner, a technical owner and a stable ID.
- [ ] Each agent authenticates with its own identity, and you have revoked one agent's credentials without affecting another.
- [ ] Each agent has a Trust Profile with an explicit list of restricted actions and a human-approval rule.
- [ ] A tool call outside an agent's allowed actions is denied by the runtime and logged, which you have tested.
- [ ] Every agent has a written autonomy level, with promotion and demotion criteria.
- [ ] You can trace one real outcome back through action, policy, approval, model and data.
- [ ] The kill switch has been exercised in the last quarter, and the time taken was recorded.

## Next steps

- [How to set up human approval for AI agents](https://www.swfte.com/how-to-set-up-human-approval-for-ai-agents): design the approve and escalate decisions properly
- [How to audit AI systems](https://www.swfte.com/how-to-audit-ai-systems): test that these controls hold up under an auditor's questions
- [How to monitor AI agents in production](https://www.swfte.com/how-to-monitor-ai-agents-in-production): the traces and alerts that feed the records in step 6
- [How to detect shadow AI](https://www.swfte.com/how-to-detect-shadow-ai): find the agents that are missing from your inventory

## FAQ

### What is AI agent governance?

It is the set of controls that decides what an agent may do and proves what it did: an inventory, an identity and owner per agent, written permissions, enforced policy, human approval where risk requires it, logs and regular review. It differs from content guardrails because it covers authority, not only output.

### What is the difference between guardrails and governance?

Guardrails answer whether a given input or output should be blocked. Governance answers who is acting, allowed to do what, under which policy, with which data and model, at what risk and with what oversight, and whether you can prove it. Most agent programmes need both.

### Who is accountable for an AI agent?

A named person. Record a business owner and a technical owner for every agent in the inventory. The business owner is accountable for what the agent does and the technical owner for how it runs. An agent without an owner should lose its credentials.

### What are the levels of AI agent autonomy?

There is no single standard scale. Swfte uses five: Assist, Approve, Supervise, Autonomous and Adaptive. Whatever scale you choose, assign a level to each agent, start low, and write down what evidence moves an agent up or down.

### Does the NIST AI RMF apply to AI agents?

It is a general AI risk framework, described by NIST as for voluntary use, built around four functions: Govern, Map, Measure and Manage. It does not mention your agents by name, but you can use it to structure agent controls. Mapping your controls to it is not the same as certification.

### Do I need special software to govern agents?

Not to start. An inventory, YAML profiles in a repository, an identity per agent, a permission check in the tool layer and a log store are enough for a first version. Software helps when you have many agents and need approvals, evidence and review at scale.

## How Swfte can help

Swfte describes governance as runtime control: policy that changes what an agent can do, with identity, approvals and an audit trail. The pages below explain the model used in this guide.

- [Trust and governance fabric](https://www.swfte.com/platform/governance): the runtime governance model and the thirteen facets
- [Trust Profile](https://www.swfte.com/platform/trust-profile): the profile fields used in step 3
- [Controlled autonomy](https://www.swfte.com/platform/autonomy): the five autonomy levels
- [AI agent governance](https://www.swfte.com/ai-agent-governance): Swfte's overview of the topic

The platform is designed to let an organisation enforce these controls at runtime. How much of this is available to you depends on your deployment: <governance feature availability by plan - founder to fill>. You can follow this guide with a spreadsheet and your existing identity provider.

## Sources

- [NIST: AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework): Four functions (Govern, Map, Measure, Manage), AI RMF 1.0 released 26 January 2023, voluntary use, Generative AI Profile NIST-AI-600-1 of 26 July 2024.
- [IEC Webstore: ISO/IEC 42001:2023](https://webstore.iec.ch/publication/90574): Standard title, scope as an AI management system, publication date December 2023 (read as a search summary of the standards catalogue; the iso.org page returned 403).
- [Microsoft Learn: What is Microsoft Entra Agent ID?](https://learn.microsoft.com/en-us/entra/agent-id/what-is-microsoft-entra-agent-id): Agent identities as a distinct construct, adaptive access, lifecycle, logging, and the separate licence note; page dated 2026-04-14.
- [OWASP GenAI: LLM06:2025 Excessive Agency](https://genai.owasp.org/llmrisk/llm062025-excessive-agency/): Three root causes (functionality, permissions, autonomy), human-in-the-loop for high-impact actions, authorisation in downstream systems rather than by the LLM, logging and monitoring.

Last verified against these sources on 2026-10-06.
