Why AI Governance Must Run at Runtime, Not in a Policy Document
A policy PDF cannot stop an agent. Why AI governance must run at runtime to allow, deny, escalate and prove.
AI governance has to run at runtime because an agent acts in milliseconds and a policy document does not. If the rule only exists in a PDF, a committee minute or a training slide, it changes what people intend and nothing about what the software can actually do. Governance that runs inside the AI system is different: the policy is evaluated on each request, it can allow, deny, filter or escalate the action, and it leaves a record you can hand to a reviewer. That is the idea behind governance as part of Sovereign Intelligence, and the rest of this post explains what it means in practice.
What is paperwork governance, and why does it fail for AI?
Paperwork governance is the familiar pattern: write an AI policy, publish it, run an annual attestation, and review incidents after they happen. It works tolerably for slow processes where a human is in every loop. It fails for AI systems for three reasons.
First, the unit of action is too small and too fast. An agent that reads a supplier file, calls a pricing tool and drafts a purchase order does dozens of things in a minute. No reviewer sees them as they happen, so a rule that depends on a person remembering it is a rule that depends on luck.
Second, the system's behaviour is not fully specified in advance. A traditional application does what its code says. An agent chooses which tool to call and with which arguments, based on a model's output. You cannot audit the code and conclude the behaviour is safe, because the behaviour is decided at run time.
Third, the evidence is thin. When governance lives in documents, the proof that it was followed is usually a signed form, not a record of what the system did. After an incident the honest answer to "which policy applied to that action?" is often "we are not sure".
None of this makes policy documents worthless. They set intent, assign accountability and define risk appetite. The point is that they are the input to governance, not the mechanism. The mechanism has to sit in the path of the action. The longer treatment is in runtime governance.
What does it mean for governance to run inside AI?
It means that every consequential step an AI system takes passes through a decision point that asks a set of questions and returns an answer before the step proceeds. The questions are the ones a good reviewer would ask if they were watching:
- Who is acting, and as which identity?
- Allowed to do what, on which system?
- Under which policy?
- With which data, and at what classification?
- Using which model?
- At what risk level?
- With what human oversight?
- And can we prove it afterwards?
The decision point returns one of a small number of verbs. The platform's governance model uses six: Allow, Deny, Warn, Filter, Escalate and Require human approval. The table shows how each changes what the agent can actually do.
| Verb | What happens to the action | Example |
|---|---|---|
| Allow | It proceeds and is recorded | Read an approved supplier record |
| Deny | It is stopped and the attempt is recorded | Read unrelated employee data |
| Warn | It proceeds and the risk is flagged to the right people | Use a model that is allowed but nearing a data-class limit |
| Filter | It proceeds with sensitive content removed or masked | Mask personal data in a prompt before it reaches a model |
| Escalate | It is routed to someone with authority to decide | A contract change outside the agent's remit |
| Require human approval | It waits until a named approver says yes | A purchase above a threshold |
The important property is that the verb is enforced, not advised. A Deny that the agent can ignore is a log line. A Deny that removes the credential or blocks the tool call changes behaviour. For the architecture behind that, see the governance layer on the platform and the reference architecture.
How are guardrails different from governance?
Guardrails answer one question: should this be blocked? They typically look at a prompt or a response and apply a filter, such as a content classifier, a topic restriction or a pattern match. That is useful and you probably want it. It is also narrow.
Governance answers a wider question: who is acting, allowed to do what, under which policy, with which data, using which model, at what risk, with what oversight, and can we prove it? A guardrail can tell you a response contained a phone number. It cannot tell you whether this agent was entitled to retrieve that customer record in the first place, whether a human should have approved the refund it triggered, or which policy version was in force last Tuesday.
A useful test: if you remove the guardrail, does the system lose a filter, or does it lose its authority model? In a governed system, removing a guardrail makes outputs riskier. Removing governance makes the system unaccountable. We go through the distinction in more depth in governance versus guardrails.
What does runtime governance look like for a real agent?
Take the AI Procurement Agent, the worked example used across the platform pages. The agent can read approved supplier information, analyse contracts, compare pricing, prepare purchase recommendations and create draft purchase orders. It cannot access unrelated employee data, approve its own high-value transaction, make payments or modify restricted records. Purchases above a threshold, contractual changes and sensitive external communications require approval.
Written as a policy sentence, that is a paragraph in a document. Enforced at runtime, it becomes a set of decisions on individual calls:
- The agent authenticates as its own identity, not as a shared service account.
- A request to read the supplier table is evaluated against the agent's allowed actions and returns Allow.
- A request to read the HR directory is evaluated and returns Deny. The attempt is recorded.
- A draft purchase order above the configured threshold returns Require human approval and waits for a named approver.
- A request to call the payments API returns Deny, because making payments is a restricted action.
- Every decision is written to a trace: identity, data accessed, model used, output, tools called, policy applied, decision, approval, action and outcome.
That last step matters as much as the enforcement. The record is what turns a behaviour into evidence. The agent's allowed and restricted actions live in a Trust Profile, a single record per AI system covering identity, owner, risk level, approved models, data classification, data residency, permitted systems, allowed and restricted actions, the human approval rule, retention, audit and policy set.
Where does the enforcement point sit?
There are three common placements, and mature setups use more than one.
| Placement | What it sees | Strength | Weakness |
|---|---|---|---|
| In the model gateway | Prompts, responses, model choice | Covers every model call, central | Does not see tool side effects |
| In the tool or action layer | Tool calls, arguments, credentials | Stops the action itself | Needs every tool to be routed through it |
| In the workflow engine | Step transitions, approvals | Natural place for human approval | Only covers governed workflows |
An agent that can reach a system through a path that bypasses the enforcement point is not governed on that path. This is the most common gap in practice, and it is why the first piece of architecture work is usually an inventory of how agents reach tools, covered in the AI estate inventory post.
Another design question is failure mode. If the policy engine is unreachable, does the action proceed or stop? For consequential actions the safe default is to fail closed, and for low-risk read-only paths an organisation may choose to fail open with a warning. That is a decision to make deliberately and write down, not one to discover during an outage.
What should a runtime record contain?
A record is useful when it lets someone reconstruct a decision without asking the engineer who built the system. At minimum:
- the identity that acted and on whose behalf
- the inputs the system saw, or a reference to them, subject to your retention and privacy rules
- the model and version used
- each tool called, with arguments and result status
- the policy and policy version that applied and the verb returned
- any approval, with approver identity and time
- the outcome
This is the chain data to model to agent to decision to action to outcome, which the platform calls traceability. For high-risk systems the EU AI Act asks for something close to it. Article 12 requires that high-risk AI systems technically allow automatic recording of events over their lifetime, and Article 26 asks deployers to keep the logs under their control for a period appropriate to the purpose, at least six months unless other law says otherwise. Those details come from secondary summaries such as the record-keeping guide at Practical AI Act and this summary of the six-month rule, so check the Official Journal text for anything you rely on in a compliance decision.
The timing has also moved. The Digital Omnibus on AI, published in July 2026, deferred stand-alone high-risk obligations from 2 August 2026 to 2 December 2027, with AI embedded in regulated products moving to 2 August 2028, according to Gibson Dunn's summary. Obligations for general-purpose AI models and the main transparency duties were not part of that delay. The sensible reading is that the delay buys time to build runtime records, not a reason to skip them. Whether any of this applies to a given system depends on its use case, jurisdiction and configuration, and counsel should decide that, not a blog post.
What about observability standards?
Traces are only as portable as the schema behind them. OpenTelemetry has published semantic conventions for generative AI, including agent spans with operations such as invoke_agent and execute_tool. They are still at Development status, not Stable, and they moved to a dedicated repository in 2026, as described in the agent spans specification. The practical advice is to emit traces in that shape where you can, keep an adapter layer so attribute renames do not break your dashboards, and treat the governance record as a separate, richer artefact. A trace says what happened. A governance record says what happened, which policy decided it, and who approved it. See also agent observability for the monitoring side.
How do you introduce runtime governance without stopping delivery?
You do not need to enforce everything on day one. A workable order:
- Observe. Put the enforcement point in the path in monitor-only mode. Record what policy would have returned. This shows you the real behaviour of your agents, which is often a surprise.
- Warn and filter. Turn on Warn and Filter for the clearest cases, such as personal data in prompts to a model that is not approved for it.
- Deny the obvious. Block actions that no one disputes: payments from an agent that is meant to draft, access to unrelated data classes.
- Add approvals. Put Require human approval on actions above a threshold and measure how often approvers change the proposal. That edit rate is evidence for later raising or lowering autonomy, as described in the controlled autonomy rollout plan.
- Close the bypasses. Remove direct credentials that let agents skip the enforcement point.
The step-by-step version of this, with the identity, policy engine, audit and evidence work, is in step 7 of the build guide. For the wider picture of why this matters to different buyers, the CISO page and the legal and compliance page list the questions to put to any vendor.
What does Swfte provide here?
Swfte's position is that governance belongs inside the platform, not beside it. Nexus captures agent actions, enforces policy in flight and traces agents and connections. Studio builds agents and workflows with approval steps. The platform provides the technical controls, governance mechanisms and evidence required to deploy AI within an organisation's applicable regulatory, security and policy requirements. The exact posture depends on the customer's use case, jurisdiction, deployment and configuration. For what is true today and what is still in progress, read the trust centre.
Where to go next
If you want to know how ready your organisation is to run governance at runtime, take the readiness self-assessment. It runs in your browser and sends nothing anywhere. If you want to talk through a specific system, talk to our team. For the underlying definitions, start at the Sovereign Intelligence pillar and the glossary.