Governance · How it works

Runtime governance: how policy changes what an AI agent can actually do

Updated 2026-10-06 · 9 min read

Short answer:Runtime governance means policy is enforced at the moment an AI agent acts, at the points where it calls a tool, reads data or invokes a model, rather than described in a document. A policy engine decides, an enforcement point applies the decision, and every step is recorded as evidence.

On this page

What does "governance happens inside AI" mean?

Most governance programmes start as paperwork: a policy, a register, a committee. Those are necessary. They do not by themselves change anything an agent does at three in the morning. Runtime governance is the part that does. A policy is expressed in a form a machine can evaluate, and that evaluation sits in the path of the agent's actions, so the policy changes what the agent can actually do.

This is the central claim of the Sovereign Intelligence Platform: governance runs through all six layers, not beside them. This article explains the mechanics. For the conceptual contrast with content filtering, read governance vs guardrails first.

What are a policy decision point and a policy enforcement point?

Access-control systems have long separated two roles, and the split works well for agents.

  • Policy enforcement point (PEP): the component sitting in the path of an action. It intercepts the request, asks for a decision, and then applies it: let the call through, stop it, mask part of it, hold it for approval.
  • Policy decision point (PDP): the component that evaluates the request against the rules and returns a decision. It does not touch the traffic.
  • Policy information: the facts the decision needs: the agent's identity, its Trust Profile, the data classification, the risk level, the user on whose behalf it acts.

Keeping the decision separate from the enforcement has practical benefits. Rules can change without redeploying every agent. Several enforcement points, at a gateway, a tool broker and a data connector, can share one set of rules. And the decision, with its inputs, can be logged in one place, which is what turns enforcement into evidence.

Where in an agent's path do you intercept?

An agent interacts with the world at a small number of boundaries. Enforcement belongs at each of them. Enforcing only at the prompt leaves the others open.

BoundaryWhat to decideTypical verbs
Model callIs this model approved for this system and this data class? Should sensitive fields be masked first?Allow, Deny, Filter, Warn
Data and context accessMay this identity read this source? Does classification or residency forbid it?Allow, Deny, Filter
Tool callMay this agent invoke this tool with these arguments? Is the action within its allowed list?Allow, Deny, Escalate, Require human approval
Agent-to-agent handoffDoes the receiving agent have authority for what is being delegated?Allow, Deny, Escalate
Output and external communicationIs this message leaving the organisation, and is it sensitive?Filter, Require human approval

The tool boundary deserves particular attention, because that is where an agent turns a decision into an effect. Open protocols help here. The Model Context Protocol gives agents a common way to reach tools and data, and it now sits under neutral stewardship in the Linux Foundation's Agentic AI Foundation, announced on 9 December 2025 (opens in a new tab). A common tool protocol gives you a single, standard place to put a PEP. See the MCP interoperability post and the hybrid integration layer write-up for the integration side.

Whose identity does an agent act under?

Every policy decision starts with "who is acting". For agents there are two answers, and conflating them is a frequent source of over-permissioned systems.

  • The agent's own identity: a service identity with a defined owner, a Trust Profile and a narrow set of permissions. It exists whether or not a person is involved.
  • The on-behalf-of user: when an agent acts for a person, the person's entitlements should limit what it can reach. An agent acting for a junior analyst must not read what the analyst cannot.

A sound default is the intersection: the agent can do only what both its own profile and the requesting user are allowed to do. Credentials should be short-lived, scoped to the tool and never embedded in prompts. Log both identities on every record. The CISO page lists the questions to put to any vendor about this, and step 5 of the build guide covers implementation.

How does the Trust Profile feed policy?

A policy engine needs facts about the thing it is governing. The Trust Profile is that record for each AI system. Its fields are: identity, owner, risk level, approved models, data classification, data residency, permitted systems, allowed actions, restricted actions, human approval rule, retention, audit and policy set.

Here is the worked example from the platform, reused verbatim:

Read it as inputs to decisions. A call to a model not in "Approved models" is denied. A read from a system not in "Permitted systems" is denied. A request to refund is a restricted action, and where the profile says approval is required above a threshold, the enforcement point holds the action and creates an approval task. The profile turns a description of an agent into something enforceable. More detail is on the Trust Profile page.

Should enforcement fail closed or fail open?

When the policy decision point is unreachable, or a rule errors, the enforcement point has to choose. Fail closed means the action is denied. Fail open means it proceeds. Neither is right everywhere.

  • Fail closed for actions with external effects, writes to systems of record, spending, and anything touching restricted data. The cost is availability. Plan for that with redundancy and clear operator alerts.
  • Fail open, with a loud warning and full logging, only for low-risk, read-only, easily reversible paths where a pause would cost more than the exposure. Make that a recorded decision per system, not a default nobody chose.
  • Degrade to a human: when in doubt, route to an approver instead of silently proceeding or silently failing.

Latency

Every policy check sits on the path of an action, so latency is a design constraint. The budget depends on the workflow. An agent that works for minutes on a research task tolerates a different overhead than a voice assistant mid-conversation. Measure decision latency in your own environment rather than relying on a quoted figure. Practical levers include evaluating rules locally next to the enforcement point, caching identity and profile lookups with short lifetimes, and keeping policies expressed as simple conditions on known fields.

Other failure modes to design for

  • Bypass paths: a direct connector or credential that skips the enforcement point. Inventory every route an agent has to a system.
  • Stale policy: an agent running on cached rules after a change. Version policies and record which version decided each action.
  • Approval fatigue: so many approval requests that reviewers click through. Tune thresholds from the record, and aggregate where it is safe.
  • Silent drift: model or tool changes that alter behaviour without any policy change. Monitor outcomes, not only decisions.

How do approval queues and policy change control work?

When a policy returns "Require human approval", the action is paused and a task goes to a named approver, with the context needed to decide: who or what is acting, what it wants to do, which data it used and why the rule fired. The approver's decision, identity and time are recorded with the action. Escalation rules need a fallback so a request does not sit unseen if the approver is away.

Policies themselves need change control, because a policy edit changes what every governed agent can do.

  1. Keep policies as versioned artefacts with an owner and a change history.
  2. Review changes the way you review code, with a second person for anything that widens authority.
  3. Test a change against recorded scenarios before it goes live, including cases that must be denied.
  4. Roll out gradually where possible: warn-only first, then enforce.
  5. Record the policy version on every decision, so a past action can be explained under the rules of the day.

Agents should not be able to modify the policies that bind them. Level changes in controlled autonomy are made by the accountable owner, not by the agent.

What does the audit record contain?

Enforcement without a record is hard to trust. The record is what turns a decision into evidence. For each governed action, capture enough to reconstruct the chain from data to outcome. The AI Procurement Agent example lists the fields: identity, data accessed, model used, output, tools called, policy applied, decision, approval, action and outcome.

Add the policy version, the on-behalf-of user, timestamps and a correlation identifier that ties the steps of one task together. Store it where you control it and can export it, for the reasons set out in governance sovereignty. The monitoring and audit automation workflow shows how teams put such records to work.

Traces and standards

Observability tooling is converging on shared vocabulary. The OpenTelemetry GenAI semantic conventions define span types for agent work, including create_agent, invoke_agent and execute_tool operations. They are still at Development status, not Stable, and have moved to a dedicated repository, so treat attribute names as subject to change (agent spans specification (opens in a new tab)). Emitting traces in that shape is sensible, as long as your governance record does not depend on the attribute names staying put. See also agent observability.

Logging is also where several regulatory duties meet engineering. The EU AI Act requires high-risk systems to support automatic event logging, and places log-keeping duties on deployers. Dates for high-risk obligations moved under the Digital Omnibus on AI, Regulation (EU) 2026/1744, which defers stand-alone Annex III obligations to 2 December 2027 (Gibson Dunn summary (opens in a new tab)). Check the Official Journal text for your own obligations. Swfte provides the technical controls, governance mechanisms and evidence for deployment within your applicable requirements. The exact posture depends on your use case, jurisdiction, deployment and configuration.

How do you test policies before trusting them?

A policy you have not tested is a hope. Test them the way you test software, with cases that must pass and cases that must fail.

  • Positive cases: the agent does its normal job and is allowed. If legitimate work is blocked, people will route around the system.
  • Negative cases: every "cannot" in the profile has a test that tries it and expects a denial. For the Procurement Agent, that includes attempting a payment and attempting to approve its own high-value transaction.
  • Approval cases: a request above threshold must pause, route to the right approver, and resume only on approval.
  • Replay: run recorded real requests against a new policy version and diff the decisions before release.
  • Adversarial cases: tool results and documents that try to talk the agent into exceeding its authority. The permission layer should hold even if the model is persuaded.
  • Failure injection: take the decision point offline and confirm the system fails the way you decided it should.

Keep these tests in the same repository as the policies and run them on every change. Then watch production: decisions per verb, approval rates and times, denials by rule, and cases where an agent repeatedly hits the same limit. Those signals tell you whether to tighten, relax or redesign. The full lifecycle appears in step 7 of the build guide, and the end-to-end layout in the architecture reference. The practical case for doing this early is made in why AI governance must run at runtime.

In the platform, Nexus is the entry point for capturing agent actions, enforcing policy in-flight and tracing every agent and connection, within the wider governance layer.

Frequently asked questions

What is runtime governance for AI?

It is enforcement of policy at the moment an AI system acts, at the model, data, tool and handoff boundaries, with every decision recorded. It differs from documentation-only governance because the policy changes what an agent can actually do.

What is the difference between a policy decision point and an enforcement point?

The decision point evaluates a request against policy and returns a decision. The enforcement point sits in the path of the action, asks for that decision and applies it by allowing, denying, masking, escalating or holding for approval.

Should an agent run with its own identity or the user's?

Both should be known. A sound default is the intersection: the agent may do only what both its own profile and the requesting user are allowed to do. Log both identities on every action.

What happens if the policy engine is unavailable?

That is a design decision made per system. Actions with external effects or restricted data should fail closed. Only low-risk, read-only, reversible paths should fail open, with loud warnings and full logging. Routing to a human is a third option.

Are OpenTelemetry GenAI conventions stable?

Not yet. As of mid-2026 the GenAI and agent span conventions are at Development status and have moved to a dedicated repository. Use them for traces, but do not make your audit record depend on attribute names that may change.

Does runtime governance work with MCP tools?

Yes. A shared tool protocol gives you a single, standard boundary at which to place an enforcement point, so tool calls can be allowed, denied or sent for approval consistently across agents.

Sources cited

Put runtime governance into practice

Start with one entry point. Add intelligence, agents, workflows and infrastructure as you prove value. Or read the step-by-step build guide and take the readiness assessment.

Ready to build with Swfte?

One platform for the agents, models and workflows your team ships. Free to start, no card required.