← The journal
Guides

A Sovereign AI Reference Architecture: Planes, Boundaries and the Life of an Agent Action

A reference architecture for sovereign AI: control and data planes, trust boundaries and one agent action traced.

Swfte Journal / Guides

A sovereign AI reference architecture is a way of drawing your AI estate so that you can answer, for every component, who controls it, where it runs and what it is allowed to touch. The practical shape is two planes (a control plane that decides and a data plane that does the work), a small number of trust boundaries, and a single enforcement path that every agent action travels through. This post is the practitioner's walk-through. The component-by-component catalogue lives in the architecture reference deep dive, and the build sequence is in the step-by-step guide. Here we focus on how the pieces behave together.

Why draw the architecture as planes and boundaries?

Because most sovereignty failures are not failures of one component. They are failures at a join. A model is hosted exactly where policy says, but the prompt log is shipped to a third-party analytics service. An agent is properly scoped, but its credentials are a shared key held in a workflow tool. Drawing planes and boundaries forces you to ask what crosses each line and under whose control.

The same view helps with the broader meaning of Sovereign Intelligence: the organisation's ability to retain meaningful control over its AI estate, which is about control, not just where a server sits. A diagram that only shows server locations misses most of that.

What are the two planes?

The control plane holds everything that decides what is allowed. It contains identity, the policy engine, the Trust Profile store, the model and tool registry, approval routing, and the audit and evidence store. It is small, changes slowly and is the part you most want to own. Compromise or loss of the control plane is the worst case for sovereignty because it is where authority comes from.

The data plane holds everything that does the work. It contains inference serving, retrieval and vector stores, agent runtimes, workflow engines, tool connectors and the business systems agents touch. It is large, changes quickly and handles the sensitive data itself.

The two planes talk through narrow, well-defined interfaces. The data plane asks the control plane for decisions and sends it records. The control plane pushes policy and registry state down. Keeping that interface narrow is what lets you place the planes differently. For example, you can run inference on a dedicated cluster in your own facility while the control plane is managed for you, or the reverse.

PlaneContainsChangesSovereignty question
ControlIdentity, policy engine, Trust Profiles, registries, approvals, audit and evidenceSlowlyWho can change a rule or read the evidence?
DataInference, retrieval, agent runtimes, workflows, connectorsOftenWhere does data go and who can reach it?

Where do the trust boundaries sit?

Mark five boundaries on the diagram and write next to each one what is allowed to cross it.

  1. Organisation boundary. Between your estate and the outside world: model providers, SaaS tools, partner systems. Egress here is the classic data-sovereignty concern, covered in data sovereignty.
  2. Environment boundary. Between production, staging and development. Agents and prompts that work in a sandbox should not carry sandbox credentials or data into production.
  3. Data-classification boundary. Between data classes, for example public, internal, confidential and restricted. A model approved for internal data is not automatically approved for confidential data.
  4. Agent boundary. Between an agent's identity and everyone else's. Each agent should act under its own identity with its own permissions.
  5. Control-plane boundary. Between the policy engine and everything it governs. Nothing the data plane does should be able to rewrite policy or edit its own audit records.

When you write down what may cross each boundary, you will often find at least one that is currently "anything". That is the first piece of work, and the AI estate inventory post describes how to find them.

What happens to a single agent action, end to end?

Follow one action through the system. Take an agent that, on request, drafts a purchase order. The numbered stages are the enforcement path.

  1. Request arrives. A user or a workflow asks the agent to prepare a purchase recommendation. The request carries the user's identity and the workflow's identity.
  2. Agent identity is resolved. The runtime authenticates the agent as its own principal and loads its Trust Profile: owner, risk level, approved models, data classification, permitted systems, allowed and restricted actions, human approval rule, retention and policy set.
  3. Context is retrieved. The agent asks the retrieval layer for supplier records. The retrieval layer applies the data-class rules and the permissions of the requesting identity, so the agent sees only what that user could see, further narrowed by the agent's own scope.
  4. Model call is routed. The prompt goes to the model gateway. The gateway checks the approved-model list, the data class of the context and the residency rule, and picks an allowed model. A filter may mask personal data on the way out. The model routing post covers cost and failover on the same path.
  5. Tool call is proposed. The model proposes calling the pricing tool and then a draft-purchase-order tool. Each proposal is sent to the policy engine before it executes.
  6. Policy decides. For each call the policy engine evaluates identity, action, data class, model, risk and oversight rules, and returns Allow, Deny, Warn, Filter, Escalate or Require human approval. Reading a price list returns Allow. A draft above the approval threshold returns Require human approval.
  7. Approval is routed, if needed. A named approver receives the draft with the context and the reasons. Their decision, identity and time are recorded.
  8. Action executes. The tool connector runs the call using a scoped, short-lived credential rather than a standing secret.
  9. Record is written. The control plane stores the full trace: identity, data accessed, model used, output, tools called, policy applied, decision, approval, action and outcome.
  10. Outcome feeds back. Whether the approver accepted, edited or rejected the draft is stored, and becomes evidence for the next autonomy decision.

The point of the walk-through is that stages 5 and 6 are in the path. If a tool can be reached without them, the architecture has a bypass, and the right fix is architectural, not a stern note in a policy.

How should you place these components in practice?

There are four common topologies. They are not tiers or maturity levels. They are different answers to different constraints.

TopologyControl planeData planeTypical reason
Managed cloudVendor-operatedVendor-operated, regionalFastest start, lowest operational load
Dedicated environmentVendor-operated or sharedIsolated VPC or bare metal in your cloud account or a provider'sIsolation, residency choice, predictable capacity
Private or on-premiseYoursYour data centre or facilityStrict control, air-gap or critical-infrastructure rules
HybridSplit by sensitivitySensitive workloads local, elastic workloads in cloudMixed data classes, burst capacity

On Swfte, customer data is held in AWS eu-west-1 (Ireland) today. Private, dedicated and hybrid deployment, and data residency in other locations, are what the platform is designed for and are scoped with you through a dedicated deployment engagement rather than being self-serve. See infrastructure and dedicated cloud, and the trust centre for the current status of each item. The economics of running your own inference are discussed in private AI on-premise economics.

One design rule holds in every topology: the control plane should be portable. If your policies, Trust Profiles and evidence can only live in one vendor's store, you have built dependency into the part you most wanted to control. This is the substance of supply-chain sovereignty, and the EU Data Act's switching rules point the same way for cloud services in general. The Act has applied since 12 September 2025 and prohibits switching charges from 12 January 2027, according to this explainer on Chapter VI.

What belongs in the data plane's inference tier?

Inference deserves its own note because it is where cost and sovereignty meet. Serving large models efficiently is largely a memory-management problem. The vLLM paper by Kwon and colleagues, presented at SOSP 2023, introduced PagedAttention, which manages the key-value cache in paged blocks to cut wasted memory, and continuous batching is a separate scheduler-level technique (paper). For architecture purposes the consequences are practical. Choose a serving engine that supports these techniques, plan GPU capacity around your longest realistic context rather than the average, and keep the engine behind your gateway so you can swap it. Our vLLM deep dive goes through the tuning, and the GPU reference covers hardware.

How do context, memory and retrieval fit?

Retrieval and memory are part of the data plane but they carry the organisation's most valuable material. Three rules keep them sovereign.

  • Permissions travel with the chunk. Index source permissions along with content, and filter at query time, so a retrieval hit never reveals something the requester cannot open. The RAG architecture guide describes the patterns.
  • Lineage is stored. Each chunk records where it came from and when, which is what makes later evidence possible.
  • Memory is owned. What agents remember and derive is intelligence sovereignty territory. You should be able to export it, delete it and inspect it.

What are the common mistakes?

  • One shared credential for all agents. It makes per-agent identity impossible and turns every incident into a guessing game.
  • Policy in the application code of each agent. It drifts. Put policy in the control plane and call it.
  • Evidence in the same store the agents can write to. An agent, or an attacker driving it, should not be able to alter its own record.
  • Treating the gateway as the whole answer. The gateway sees model calls, not side effects. Tool and workflow paths need enforcement too, as covered in why governance must run at runtime.
  • Designing for the final topology on day one. Start with the narrow interface between the planes and the rest can move.

How does this map to the six layers?

The six platform layers map onto the planes cleanly. Sovereign Infrastructure, Data and Context, and Intelligence and Models are mostly data plane. Governed Agents and Governed Workflows sit across both, because the runtime is in the data plane but the authority comes from the control plane. AI Solutions sit on top. The trust and governance fabric is not a seventh layer. It is the control plane's reach into all six. The platform overview shows the layers and the layer-by-layer build post walks through them in order.

Where to go next

Draw your own version: two planes, five boundaries, one enforcement path, and mark each crossing with who controls it. Gaps will show up quickly. To score where you stand and get a suggested starting point, use the readiness self-assessment, which runs entirely in your browser. For help mapping this to your estate, talk to our team, and for the definitions behind the terms, see the glossary.

Related: Swfte Connect is the model gateway, designed to run in your own cloud or data centre; see self-deploying Connect.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Ready to build with Swfte?

One platform for the agents, models and workflows your team ships. Free to start, no card required.