Step-by-step guide

How to Build a Sovereign Intelligence Platform

Updated 2026-10-06 · 39 min read · 10 steps

Short answer:To build a Sovereign Intelligence Platform, work through ten steps in order: define your sovereignty requirements, choose infrastructure, prepare data and context, select models, build governed agents and workflows, wire in the trust fabric, set autonomy levels, deliver solutions, and operate and learn. Each step has decisions, pitfalls and a checklist.

The ten steps

Before you start

A Sovereign Intelligence Platform is the set of infrastructure, data, models, agents, workflows and controls that lets an organisation run AI while keeping meaningful control over it. The category pillar defines the idea. This guide is the practical companion: the order of work, the decisions you will face at each stage, and the mistakes that cost teams the most.

The ten steps follow the six layers of the platform, with the trust and governance fabric treated as a thread that runs through all of them rather than a final step. You will see governance decisions in step 1, in step 2 and in step 5, not only in step 7. That is deliberate. Controls added at the end are paperwork. Controls designed in from the start change what the AI can actually do.

Who this is for

Platform architects, heads of AI and data, security and risk owners, and technical founders who have moved past a first demo and now need AI that survives a security review, an audit and a change of vendor. If you are still deciding whether the investment is justified, take the readiness assessment first.

How to use this guide

Read it once end to end, then return to individual steps as you reach them. Every step has the same shape: goals, decisions, architecture notes, pitfalls, a checklist, and how Swfte approaches the step. Nobody should try to finish all ten at once. The land-and-expand ladder is the realistic path: land small, prove value, then add intelligence, agents, workflows and infrastructure. Where a step says "designed for", it describes what the platform is built to let you do, and your exact posture depends on your use case, jurisdiction, deployment and configuration.

Step 1 of 10

01Define your sovereignty requirements

Decide what control means for your organisation before choosing any technology, so every later choice has a test to pass.

Goals

  • Write down, per kind of sovereignty, what you must control and what you are content to delegate.
  • Classify the data, workloads and decisions the AI will touch, and rank them by sensitivity and consequence.
  • Map the laws, contracts and internal policies that constrain where AI may run and what it may do.
  • Produce a short requirements document that architecture, security and procurement all sign.

Decisions you will face

Which of the seven sovereignties are hard requirements?
Data, infrastructure, model, intelligence, operational, governance and supply-chain sovereignty each carry a different cost. Rank them as must, should or can delegate. Most organisations find that two or three are must-haves and the rest are negotiable.
How will you classify data for AI use?
Keep it to three or four classes, for example public, internal, confidential and restricted. For each class state which locations may process it: on the device, on your own infrastructure, in a dedicated cloud environment, or through a hosted model API.
Which jurisdictions and regulations apply to each workload?
List the legal regimes per workload rather than per company: data protection, sector rules, public-procurement rules and the EU AI Act where relevant. A single organisation often needs different answers for different workloads.
Who is accountable for each AI system?
Name an owner per system and an executive sponsor per domain. Accountability without a name is the most common gap found in AI audits.
What is your exit position?
Decide how quickly you must be able to replace a model, a cloud or a vendor. This single number shapes the architecture more than any feature list.

Architecture notes

Start from the definition: sovereignty is the organisation's ability to retain meaningful control over its AI estate. It is about control, not just where a server sits. A server in your own country that you cannot inspect, patch or leave is a location, not sovereignty. The deep dive on sovereign AI versus sovereign intelligence explains why location is only one input.

Jurisdiction matters because legal exposure follows the operator, not only the data centre. The US CLOUD Act, for instance, lets US authorities compel disclosure from US-based providers for data in their possession, custody or control wherever it is stored, which sits uneasily with GDPR Article 48 according to CSIS analysis (opens in a new tab). That is a reason to ask who operates a service, not only where it runs.

The European Commission's Cloud Sovereignty Framework (opens in a new tab) gives a useful vocabulary even if you are not buying public-sector cloud. It scores services against eight objectives (strategic, legal and jurisdictional, data and AI, operational, supply chain, technology, security and compliance, and environmental) on a five-level scale. You can borrow the idea of setting a minimum level per objective instead of one blended score.

A requirements matrix. The entries are illustrative; fill yours in per workload.
SovereigntyQuestion to answerTypical evidence
DataWhere may each data class be processed, and who can access it?Classification policy, residency map, access reviews
InfrastructureWhere does AI run and what does it depend on?Deployment diagram, dependency inventory
ModelWhich models are approved, and who can change them?Approved-model list, licence review, eval results
IntelligenceWho owns the prompts, memory and derived insight?Contract clauses, export test
OperationalWhat may agents do on your behalf?Trust Profiles, permission scopes
GovernanceWhere do the policies, oversight and proof live?Policy repository, audit trail, evidence packs
Supply chainWhich suppliers can you replace, and how fast?Supplier register, tested exit plan

Finish with a one-page requirements statement. It should be short enough that a board member reads it and specific enough that an engineer can fail a design against it. For each workload, record the data class, the permitted processing locations, the approved model families, the maximum autonomy level and the named owner. The later steps in this guide each produce a piece that must match this page.

The seven sovereignty deep dives cover each kind in more depth: infrastructure, model, intelligence, operational, governance and supply-chain sovereignty.

Common pitfalls

  • Treating residency as the whole requirement. Data in the right country can still be reachable through a provider's home jurisdiction, a subprocessor or an admin console.
  • Writing requirements per technology (we need Kubernetes) instead of per outcome (we must be able to move any workload within a quarter).
  • Letting one team own the document. Security, legal, data, engineering and the business each see a different risk; the document fails if any of them are missing.
  • Skipping the exit position, then discovering at renewal that the model, the vector store and the prompts cannot leave.
  • Setting requirements so strict that nothing can ship, which pushes teams back to shadow AI. See the shadow AI analysis.

Checklist

  • Seven sovereignties ranked as must, should or can delegate.
  • Data classes defined, each mapped to permitted processing locations.
  • Applicable laws, contracts and policies listed per workload.
  • A named owner and executive sponsor for each planned AI system.
  • An exit position stated in time, for model, cloud and vendor separately.
  • One-page requirements statement signed by security, legal, data and engineering.
  • A review date, because requirements move when regulation and suppliers do.

Step 2 of 10

02Choose your infrastructure

Pick where AI runs, size the GPUs honestly, and choose an inference stack you can operate and replace.

Goals

  • Choose a deployment pattern per workload: public cloud, private or dedicated cloud, on-premises or hybrid.
  • Size GPU memory for model weights and KV cache, not for weights alone.
  • Select an inference serving stack that gives you throughput, isolation and observability.
  • Define the network, key-management and update path, including any air-gapped requirement.

Decisions you will face

Which deployment pattern fits each workload?
Public cloud for low-sensitivity and bursty work, dedicated or private cloud for sensitive workloads that need isolation, on-premises for the strictest data or latency needs, and hybrid when sensitivity differs across workloads. Decide per workload, not per company.
Do you rent GPUs, reserve them or buy them?
Renting suits experiments and uncertain demand. Reserved capacity suits steady inference. Owning hardware only pays back when utilisation is high and you have people to run it. Model your own utilisation before deciding.
Which inference engine will serve models?
Open-source servers such as vLLM are common for self-hosting; managed endpoints reduce operations at the cost of control. Whatever you choose, require an OpenAI-compatible or otherwise documented interface so you can swap engines.
Where do encryption keys live?
Provider-managed keys are simplest. Customer-managed keys give you the ability to revoke access. For the strictest data, keys held in your own hardware security module are worth the operational cost.
Do you need an air-gapped environment?
If yes, plan the update path first: how model weights, container images, dependencies and policy updates enter the environment, and who verifies them.

Architecture notes

Deployment patterns compared. Use as a starting point for your own assessment.
PatternControlOperational loadTypical fit
Public cloudLowest: shared tenancy, provider operatesLowestLow-sensitivity, bursty or experimental work
Private or dedicated cloudHigher: isolated environment, you set boundariesMediumSensitive workloads that need isolation without owning hardware
On-premisesHighest: you own hardware, network and physical accessHighestStrictest data, latency or air-gap needs
HybridPer workload: sensitive work stays close, other work moves outMedium to highMixed sensitivity across workloads

Size GPUs for weights plus KV cache

GPU memory has two large consumers. The first is the model weights: parameters multiplied by bytes per parameter, so roughly 2 bytes each at 16-bit precision, 1 byte at 8-bit and 0.5 bytes at 4-bit, according to a VRAM sizing guide (opens in a new tab). The second is the KV cache, the stored attention keys and values for every token in every active request. Per token it is about 2 x layers x KV heads x head dimension x bytes per element, so it grows with context length and with concurrency.

The practical consequence is that a model which "fits" on a card at idle may not serve your real traffic. Long contexts and many simultaneous users are what exhaust memory. A hypothetical example: a model whose weights take 16 GB on a 40 GB card leaves roughly 24 GB, minus runtime overhead, for the KV cache that every concurrent request must share. Work out your target context length and concurrency, then check the number rather than guessing. Mixture-of-experts models load all experts into memory even though only some are active per token.

Inference serving

Serving engines differ mostly in how they manage that KV cache and schedule requests. The vLLM paper by Kwon and colleagues introduced PagedAttention, which stores the cache in non-contiguous blocks like operating-system paging so memory is not reserved in large contiguous slabs (arXiv 2309.06180 (opens in a new tab)). Continuous batching is a separate technique at the scheduler level, adding and removing requests at every step. Our vLLM deep dive walks through both. For a sovereign platform the important properties are operational: can you pin versions, observe queue depth and latency, isolate tenants, and replace the engine without rewriting applications?

Network, keys and air gaps

  • Network. Default-deny egress from the inference environment. Allow only named destinations, such as your own vector store and approved external tools. An AI environment that can call any URL can also leak to any URL.
  • Keys. Separate keys per environment and data class. Decide who can rotate and revoke, and test revocation before you depend on it.
  • Updates. Treat model weights as software supply chain. Pin versions, verify checksums, keep a record of where each artefact came from, and stage updates in a non-production environment.
  • Air gap. An air-gapped deployment has no internet path, so mirrors for packages, container images and weights are required, along with an offline process for policy and model updates. It also removes hosted tools and external search from the agents' toolbox, which is a feature and a limitation.
  • Portability. Keep infrastructure as code, avoid provider-specific accelerators in the application layer, and run an exit drill. The on-premises economics post covers when owning hardware pays back.

The EU Data Act adds a regulatory reason for portability. Its cloud-switching rules apply from 12 September 2025, and switching charges, including egress charges for switching, are prohibited from 12 January 2027 (summary of Chapter VI (opens in a new tab)). Check the text and your contracts for how it applies to you. See also the deep dive on infrastructure sovereignty.

Common pitfalls

  • Sizing GPUs from the weight file alone and discovering at launch that the KV cache for real concurrency does not fit.
  • Buying hardware before measuring utilisation. An idle GPU is the most expensive part of a private deployment.
  • Choosing an inference stack tied to one accelerator or one provider's API, which quietly removes the exit you defined in step 1.
  • Leaving egress open from the model environment. Prompt injection plus open egress is a data-exfiltration path.
  • Treating an air gap as a security guarantee. It reduces exposure, but insiders, supply-chain tampering and removable media remain.

Checklist

  • Deployment pattern chosen and justified per workload.
  • GPU memory budget worked out for weights, KV cache and overhead at target context length and concurrency.
  • Inference engine selected with a documented, replaceable interface.
  • Default-deny egress configured and tested from the inference environment.
  • Key ownership, rotation and revocation procedure written and exercised once.
  • Artefact provenance recorded for every model weight and container image.
  • Update path defined, including the offline path if air-gapped.
  • An exit drill scheduled: move one workload to another environment and measure the effort.

Step 3 of 10

03Prepare data and context

Give AI the context to understand the organisation, using retrieval, memory and lineage that respect who may see what.

Goals

  • Inventory the sources AI will read and decide which are in scope for the first release.
  • Build a retrieval layer that honours each source system's permissions.
  • Decide what the AI may remember, for how long, and who can inspect or delete it.
  • Record lineage so any answer can be traced back to the documents and data behind it.

Decisions you will face

Retrieval only, or fine-tuning too?
Start with retrieval-augmented generation. It keeps knowledge outside the model, lets you update it instantly, and lets you enforce access control at query time. Fine-tune only when you need a behaviour or style retrieval cannot provide.
Vector search alone, or hybrid?
Pure vector search misses exact identifiers, codes and names. Hybrid search runs keyword and vector retrieval in parallel and fuses the rankings, often with reciprocal rank fusion, then optionally reranks.
Do you need a knowledge graph?
Add one when questions depend on relationships (who owns this contract, which systems depend on this supplier) that text chunks capture poorly. Do not start there; add it when retrieval quality plateaus.
Which memory types will agents have?
Working memory for the current task, episodic memory of past runs, and semantic memory of durable facts. Each needs a retention rule and an owner.
Where do the index and embeddings live?
Embeddings can leak the content they were made from, so treat the vector store as the same sensitivity class as the source data and keep it in the same boundary.

Architecture notes

Retrieval-augmented generation was introduced by Lewis and colleagues in 2020 (opens in a new tab) as a way to give a language model access to knowledge at inference time without retraining it. For a sovereign platform its main virtue is governance: knowledge stays in systems you control, and every answer can cite what it read. Our RAG architecture guide covers the engineering in depth; this section covers the sovereignty-specific decisions.

Permission-aware retrieval

The most common data-layer failure is a retrieval index that flattens permissions. Every chunk must carry the access-control information of its source, and the query must run as the requesting identity, or as an agent identity scoped no wider than the user's rights. Otherwise the AI becomes a way to read documents the user could not open directly. Test with two accounts, one that should see a document and one that should not, after each source is added.

Components of the data and context layer.
ComponentJobSovereignty question
Connectors and ingestionPull content from systems of record and keep it currentDoes ingestion preserve permissions and deletions?
Chunking and enrichmentSplit documents sensibly and attach metadataIs classification carried into every chunk?
EmbeddingsTurn text into vectors for semantic searchWhich embedding model, and where does it run?
Vector or hybrid indexRetrieve candidates by meaning and by keywordIs the store inside the data boundary and exportable?
Knowledge graphRepresent entities and relationshipsWho curates it, and can it be exported?
Memory storePersist what agents learn between runsRetention, inspection and deletion rules?
Lineage and logsRecord which sources produced which answerCan you reconstruct an answer a year later?

Chunking, embeddings and hybrid search

Chunk by document structure (sections, clauses, table rows) rather than a fixed character count, and keep the title and path with each chunk. Choose an embedding model you can run inside your boundary, and record its version, because changing it means re-embedding everything. For retrieval, hybrid search combines lexical scoring such as BM25 with semantic vectors. Averaging their raw scores does not work because the scales differ, so rankings are fused instead; reciprocal rank fusion scores each document by summing 1 divided by a constant plus its rank in each list, as described by retrieval practitioners (opens in a new tab). Finish with a reranker if latency allows.

Memory and lineage

Memory is where intelligence sovereignty is won or lost. What an agent remembers about your customers, suppliers and processes is derived insight, and it should be an asset you can inspect, export and delete. Store memory in your own boundary in a documented format. Record lineage for each answer: the identity, the query, the chunks retrieved with their versions, the model, the output. That record is what lets you answer a regulator, a customer or your own auditor later. See data sovereignty and intelligence sovereignty.

Common pitfalls

  • Indexing everything at once. Start with a small, well-permissioned set of sources and expand one at a time.
  • Ignoring deletions. If a document is removed or a permission is revoked at the source, the index must follow, or the AI keeps answering from it.
  • Storing embeddings or chat history in a convenient third-party service outside the data boundary.
  • Changing the embedding model without a re-embedding plan, which silently degrades retrieval.
  • Evaluating retrieval by eye. Build a small question set with known source documents and measure whether the right chunks come back.

Checklist

  • Source inventory with owner, classification and permission model for each source.
  • First-release scope limited to a few high-value sources.
  • Ingestion preserves permissions, deletions and metadata.
  • Two-account permission test passes for every connected source.
  • Hybrid retrieval configured, with a retrieval-quality test set.
  • Embedding model and version recorded; re-embedding plan written.
  • Memory types, retention rules and deletion procedure defined.
  • Lineage record captured per answer and retained according to policy.

Step 4 of 10

04Select and evaluate models

Choose which models are approved, route work between them, and prove quality with evaluations you own.

Goals

  • Define an approved-model list tied to data classes.
  • Decide where open, private and tailored models are worth their cost.
  • Build routing that picks a model per task, with failover.
  • Create evaluations that gate every model or prompt change.

Decisions you will face

Which models are approved for which data class?
Map each data class to the model deployments allowed to see it. Restricted data may only reach models you host; internal data may reach a hosted API with contractual limits; public data may go anywhere approved.
Open-weight, hosted proprietary, or both?
Open-weight models give you control of deployment and lifecycle but cost operations. Hosted proprietary models often lead on capability but bring a dependency. Most estates use both and route between them.
Is tailoring worth it?
Prompting and retrieval first, then lightweight adaptation, then fine-tuning. Each step adds an asset you must version, evaluate and retire.
How is routing decided?
By task type, data class, latency target and cost, with a policy-controlled fallback. Routing rules are policy, so they belong under change control.
What licence terms are acceptable?
Review each model licence for commercial use, field-of-use limits, redistribution and attribution before it enters the approved list.

Architecture notes

Model sovereignty means selection, deployment, customisation and lifecycle stay in your hands. In practice that means a model registry: every approved model has an entry recording its provider or source, version, licence, hosting location, permitted data classes, evaluation results and retirement date. The model sovereignty deep dive expands on each field.

Three ways to source a model.
OptionStrengthCostSovereignty note
Hosted proprietary APIOften strongest capability, no serving workDependency on one vendor's terms and availabilityCheck data handling terms and exit path; keep a fallback
Open-weight, self-hostedYou control deployment, versions and data pathGPU and operations costVerify the licence and the weights' provenance
Tailored (adapted or fine-tuned)Fits your domain, tone and tasksTraining data, evaluation and retraining effortThe tailored weights and training data become assets you own and must govern

Routing

A routing layer sits between applications and models. It selects by task and policy, retries on failure, falls back to another provider, records cost and latency, and enforces the approved-model list. Applications call one interface, which is what makes a model swap a configuration change rather than a rewrite. Our post on model routing covers the cost side, and the model exit cost audit is a useful template for testing how locked in you are.

Evaluations as the gate

Public leaderboards do not tell you whether a model works on your documents and your tasks. Build an offline evaluation set from real, de-identified examples of the work: inputs, expected outputs or scoring rubrics, and the edge cases that hurt when wrong. Run it on every candidate model, every prompt change and every retrieval change. A regression gate means a change that lowers scores on critical cases does not ship. Add safety evaluations too: prompt-injection attempts, requests that should be refused, and attempts to extract data the user may not see. The OWASP Top 10 for LLM applications (opens in a new tab) is a reasonable checklist of the failure classes to include.

  • Offline evaluation. Fixed question sets scored by rubric, by reference answer or by a calibrated judge model whose agreement with humans you have measured.
  • Online evaluation. Sampled production traffic reviewed by people, with feedback captured per answer.
  • Regression gate. A written threshold per critical scenario. Changes below it are blocked.
  • Drift watch. Providers update hosted models. Re-run the suite on a schedule so silent changes surface.

Common pitfalls

  • Picking a model from a leaderboard and skipping your own evaluations.
  • Hard-coding one provider's SDK throughout the application, so a switch means a rewrite.
  • Fine-tuning on data you are not permitted to use for training, or without recording what went in.
  • Ignoring licence terms until legal review at launch.
  • Treating a hosted model as a fixed artefact. Versions change, sometimes without notice.

Checklist

  • Model registry exists with source, version, licence, location and permitted data classes.
  • Approved-model list mapped to data classes and enforced in code, not by memo.
  • Routing layer in place with failover and cost and latency capture.
  • Offline evaluation set built from real tasks and stored under version control.
  • Regression gate defined for critical scenarios.
  • Safety evaluation covering injection, refusal and data-leak attempts.
  • Licence review recorded for every approved model.
  • Scheduled re-evaluation to catch drift in hosted models.

Step 5 of 10

05Build governed agents

Give each agent an identity, the least privilege it needs, controlled tools and a clear approval rule.

Goals

  • Give every agent its own identity, owner and Trust Profile.
  • Grant the minimum permissions for the task, and no standing access to anything else.
  • Connect tools through a controlled interface that can allow, deny or escalate each call.
  • Define what needs a human before it happens.

Decisions you will face

Does the agent act as the user, as itself, or both?
Acting as the user inherits the user's rights and attribution. Acting as itself allows narrower, auditable permissions. Many designs use both: the agent reads as the user and writes only through its own scoped identity.
How are tools exposed?
Through a gateway that authenticates the agent, checks policy per call and records it. The Model Context Protocol is a common open standard for this interface.
What requires approval?
Anything irreversible, costly, external-facing or touching restricted records. Define thresholds as policy so they can be tuned without code changes.
Single agent or several?
Prefer several narrow agents with distinct permissions over one broad agent. A narrow agent has a small blast radius and an obvious owner. See the multi-agent systems post.
Who may create agents?
Decide whether teams self-serve inside guardrails or request through review. Either way, an unregistered agent should not be able to obtain credentials.

Architecture notes

The central risk with agents is authority. OWASP's 2025 list names excessive agency, where a system grants a model too much functionality, permission or autonomy, as a distinct risk, and notes that prompt injection is often how an attacker triggers it (OWASP LLM Top 10 (opens in a new tab)). The design answer is to assume the model can be manipulated and to constrain what a manipulated agent could do. That is least privilege, enforced outside the model.

Identity and the Trust Profile

Every agent has an identity in your identity provider or a platform equivalent, an accountable owner, and a Trust Profile recording the fields below. The profile is the contract the runtime enforces. See the Trust Profile page for the full format.

  • Identity, owner and risk level
  • Approved models and data classification
  • Data residency and permitted systems
  • Allowed actions and restricted actions
  • Human approval rule, retention, audit and policy set

Tools through a controlled interface

Agents act through tools: APIs, databases, browsers, other agents. The Model Context Protocol standardises how models connect to tools and data, and Anthropic donated it to the Agentic AI Foundation, a Linux Foundation directed fund, on 9 December 2025 (MCP blog (opens in a new tab)). A standard interface helps, but the standard does not itself decide who may call what. Put a gateway in front of tool servers that authenticates the agent, applies policy to each call, redacts where needed and logs the result. Our posts on MCP as an interoperability standard and the hybrid MCP layer go further.

Worked example: the AI Procurement Agent

AI Procurement Agent
CanCannotRequires approvalRecords
Read approved supplier info; analyse contracts; compare pricing; prepare purchase recommendations; create draft POsAccess unrelated employee data; approve its own high-value transaction; make payments; modify restricted recordsPurchases above threshold; contractual changes; sensitive external commsIdentity, data accessed, model used, output, tools called, policy applied, decision, approval, action, outcome

Notice that the agent cannot approve its own high-value transaction. Separation of duties applies to software as it does to people. The permissions to read suppliers and to create drafts are separate grants, and the approval step is a different identity. The governed agents page and the operational sovereignty deep dive cover the pattern.

Common pitfalls

  • Giving agents a shared service account. Attribution is lost and one compromise exposes everything.
  • Relying on the system prompt to enforce limits. Instructions in a prompt are not controls.
  • Letting tool credentials sit in the agent's context where a prompt injection can read them.
  • Approval fatigue: so many approval requests that people click through. Tune thresholds and batch low-risk items.
  • No kill switch. Every agent needs a way to be suspended immediately by its owner.

Checklist

  • Each agent has a unique identity, a named owner and a Trust Profile.
  • Permissions granted per task and reviewed on a schedule.
  • Tools reached only through a gateway that checks policy and logs calls.
  • Credentials held by the gateway, never exposed to the model.
  • Approval rules defined as policy, with a named approver role.
  • A suspend control exists and has been tested.
  • Prompt-injection tests run against every tool the agent can call.
  • Agent registry lists every agent in the estate.

Step 6 of 10

06Design governed workflows

Embed AI in real processes with explicit steps, checkpoints, error handling and ownership.

Goals

  • Pick processes where AI adds value and the cost of an error is bounded.
  • Express each process as explicit steps with inputs, outputs, owners and approval points.
  • Handle failure, retries and rollback deliberately.
  • Make every run observable and attributable.

Decisions you will face

Which processes first?
Choose repeatable work with clear inputs and outputs, tolerable error cost, and an owner who wants it. Avoid first projects that touch irreversible financial or safety decisions.
Agent-led or workflow-led?
Use a fixed workflow where the steps are known, and let an agent decide only inside steps that need judgement. Deterministic scaffolding with AI inside is easier to govern than a free-roaming agent.
Where are the human checkpoints?
Before irreversible actions, at thresholds, and wherever a person is legally or contractually accountable for the decision.
How do runs fail safely?
Decide the behaviour on tool failure, model failure, policy denial and timeouts: retry, skip, escalate or stop. Never default to continue silently.
Who owns the workflow in production?
The business owner owns outcomes, a platform owner owns the runtime, and a change process governs edits to either.

Architecture notes

A governed workflow is a process definition plus the runtime that executes it. The definition names steps, the identity that performs each, the data each may read and write, the policy that applies, and where a human decides. The runtime enforces the definition and records what happened. Embedding AI in how the organisation operates is the point of this layer; governance makes it safe to scale.

  1. Trigger. An event, a schedule or a person starts the run, and the trigger identity is recorded.
  2. Gather context. Retrieve data and documents as the acting identity, with permissions applied.
  3. Decide or draft. A model or agent produces a recommendation or a draft within its Trust Profile.
  4. Check. Policy evaluates the proposed action: allow, deny, warn, filter, escalate or require human approval.
  5. Act. The approved action executes through a controlled tool with an idempotency key where possible.
  6. Record. Inputs, outputs, policy decisions, approvals and outcome are written to the run record.

Treat each workflow as software. Version it, test it with representative inputs and adversarial ones, review changes, and promote through environments. Keep a rollback path: if a new version misbehaves, you must be able to revert in minutes and identify which runs it affected. Design for idempotency so a retried step does not send two emails or create two purchase orders.

Workflow-level controls complement agent-level ones. An agent can be perfectly scoped and still be placed in a workflow that routes its output to the wrong place. Review the data flow end to end, which is exactly the traceability chain the trust fabric records: data, model, agent, decision, action, outcome. See the governed workflows page and our post on controlled autonomy rollout for how approval points loosen over time.

Common pitfalls

  • Automating a broken process. AI makes a poor process faster, not better.
  • Hiding human approval inside a chat message with no record. Approvals are evidence and need identity and time.
  • No idempotency, so retries duplicate side effects.
  • Editing live workflows in place with no version history.
  • Measuring only throughput. A workflow that is fast and wrong is a liability.

Checklist

  • Process documented with owner, inputs, outputs and error cost.
  • Steps assigned identities and least-privilege access.
  • Human checkpoints defined and tested.
  • Failure behaviour defined for each step type.
  • Idempotency or duplicate protection on every side-effecting step.
  • Versioning, review and rollback in place.
  • Run records searchable by workflow, version, identity and outcome.

Step 7 of 10

07Wire the trust fabric

Connect identity, policy, audit and evidence so governance runs inside AI instead of beside it.

Goals

  • Make identity and access the basis for every AI action.
  • Enforce policy at runtime and be able to change behaviour by changing policy.
  • Capture an audit and trace record detailed enough to reconstruct any decision.
  • Produce evidence packs for reviewers without a manual scramble.

Decisions you will face

Where is policy enforced?
At the points where AI touches the world: the model gateway, the tool gateway and the workflow runtime. Enforcement inside the application only is easy to bypass.
How is policy written and versioned?
As code or structured configuration under version control, reviewed like software, with each decision recording the policy version that produced it.
What exactly is logged?
Identity, data accessed, model used, output, tools called, policy applied, decision, approval, action and outcome. Decide retention per data class and be careful about personal data in the logs themselves.
What telemetry standard do you adopt?
OpenTelemetry is the common choice for traces. Its GenAI conventions are still in development, so pin versions and expect attribute names to change.
How is evidence packaged?
Decide which reviewers (internal audit, regulators, customers) need what, and generate those views from the same underlying record.

Architecture notes

The trust and governance fabric is not a seventh layer. It runs through all six, and has thirteen facets: identity, access, data controls, policy, security, privacy, compliance, risk, human oversight, auditability, traceability, evidence and monitoring. Guardrails ask whether something should be blocked. Governance asks who is acting, allowed to do what, under which policy, with which data, using which model, at what risk, with what oversight, and whether you can prove it. The guardrails versus governance deep dive draws the line, and runtime governance explains how enforcement works.

The thirteen facets and the artefact that makes each real.
FacetArtefact or mechanism
IdentityIdentity per user, agent and workflow
AccessScoped permissions per identity, tool and data source
Data controlsClassification and residency enforcement
PolicyVersioned rules evaluated at runtime
SecurityCredential vaulting, network controls, injection testing
PrivacyRedaction and handling rules for personal data
ComplianceControls mapped to your own obligations
RiskRisk levels and thresholds that decide approval
Human oversightApproval, escalation and review points
AuditabilityRecord of actions and the policy applied
TraceabilityChain from data to model to agent to decision to action to outcome
EvidenceExportable packs for reviewers
MonitoringLive view of behaviour, cost and policy decisions

The policy engine

A policy engine takes a request (who, what action, which data, which model, what risk) and returns a verdict. The verbs are allow, deny, warn, filter, escalate and require human approval. Because the engine sits in the request path, changing a policy changes what agents can actually do, which is what "governance happens inside AI" means. Keep policy decisions fast and deterministic, and fail closed for high-risk actions.

Audit, tracing and telemetry

For traces, OpenTelemetry's GenAI semantic conventions define agent and tool spans such as invoke_agent and execute_tool, but they remain at Development status and moved to a dedicated repository in 2026 (agent spans specification (opens in a new tab)). Adopt them with a thin mapping layer so renamed attributes do not break your dashboards. Our agent observability page covers the product side.

Regulation also shapes logging. For high-risk systems, Article 12 of the EU AI Act requires the system to technically allow automatic recording of events over its lifetime, and Article 26 asks deployers to keep logs under their control for at least six months unless other law requires longer, according to practitioner summaries (opens in a new tab); confirm the exact wording in the Official Journal. The Digital Omnibus on AI (Regulation (EU) 2026/1744) moved the stand-alone high-risk obligations from 2 August 2026 to 2 December 2027, per Gibson Dunn (opens in a new tab). Whether your system is high-risk depends on its use. For wider frameworks see the NIST AI RMF (opens in a new tab), whose four functions are Govern, Map, Measure and Manage, and ISO/IEC 42001 for an AI management system. See also our EU AI Act page and NIST AI RMF page.

Common pitfalls

  • Logging everything but nothing useful: no policy version, no identity, no link between the model call and the action it caused.
  • Putting personal data into logs with no retention rule.
  • Policy that exists only as a PDF. If the runtime does not read it, it does not govern.
  • Building evidence by hand before each audit. It will be late, partial and unrepeatable.
  • Failing open when the policy engine is unreachable for high-risk actions.

Checklist

  • Identity assigned to every user, agent and workflow.
  • Policy engine in the model, tool and workflow paths, failing closed for high-risk actions.
  • Policies stored as versioned artefacts with review.
  • Audit record includes identity, data, model, tools, policy, decision, approval, action and outcome.
  • Retention set per data class, with a plan for personal data in logs.
  • Trace telemetry adopted with a mapping layer for convention changes.
  • Evidence packs generated from the same record for each reviewer type.
  • Applicable regulatory dates checked against current official sources.

Step 8 of 10

08Set autonomy levels

Decide how much each AI system may do on its own, and move it up only as evidence accumulates.

Goals

  • Assign every agent and workflow an autonomy level from L1 to L5.
  • Define the evidence required to promote and the triggers that demote.
  • Make autonomy a recorded policy setting, not an informal habit.
  • Keep the owner, not the AI, in control of changing the level.

Decisions you will face

What is the starting level for a new agent?
L1 Assist or L2 Approve for nearly everything new. Higher starting levels need a documented reason.
What evidence earns promotion?
A sustained acceptance rate for its outputs, low edit and rejection rates, no policy violations, and a clean review of sampled cases. Set the thresholds yourself, per use case.
What triggers demotion?
Policy violations, accuracy regressions, incidents, a changed model or tool, or a changed risk profile. Demotion should be automatic where possible.
Can any agent stay at a low level permanently?
Yes. Many agents should stay at L2 or L3. The right level depends on risk and reversibility, not on what is technically possible.
Who approves a level change?
The accountable owner, with risk or security sign-off for L4 and above. The agent can never raise its own level.

Architecture notes

Controlled autonomy replaces the idea of "manual versus autonomous" with five levels, each with different approval, monitoring and audit rules:

The five autonomy levels.
LevelMeaning
L1 AssistAI recommends.
L2 ApproveA human approves.
L3 SuperviseAI acts within limits and is monitored.
L4 AutonomousAI works independently within strict policy and risk bounds.
L5 AdaptiveAI improves within controlled boundaries.

Autonomy is per action, not per agent. The same Procurement Agent can be at L3 for drafting purchase orders and at L1 for anything involving contract changes. As the level rises, approval moves from every action to thresholds and exceptions, monitoring moves from review to live alerting, and the audit record deepens. Hard limits are enforced by the runtime and are not something the agent can modify. L5 is the most sensitive: an agent that improves its own behaviour must have its changes versioned and gated, and its boundaries must not be self-modifiable.

A practical promotion path: run at L1 and measure how often people accept the recommendation. Move to L2 and measure edits and rejections. Move to L3 for a bounded class of low-risk actions and watch live. Only after a recorded history of clean operation, and only where limits, rollback and monitoring are proven, consider L4 for that class. Our controlled autonomy rollout post gives a longer plan, and the autonomy page shows the controls at each level. For the policy side see capability plus control.

Common pitfalls

  • Promoting on a calendar ("after one month") rather than on evidence.
  • Setting one level for the whole agent when the risk varies by action.
  • No demotion path, so a regression stays at a level it no longer deserves.
  • Letting autonomy settings live in application code that only engineers can read.
  • Declaring L4 or L5 for marketing reasons before the controls exist.

Checklist

  • Every agent and workflow has a recorded autonomy level per action class.
  • Promotion criteria written with numeric thresholds you chose.
  • Demotion triggers defined and, where possible, automatic.
  • Level changes require the owner and are logged in the Trust Profile.
  • Hard limits enforced by the runtime.
  • Rollback and kill switch tested before any L3 or above.
  • Quarterly review of levels against evidence.

Step 9 of 10

09Deliver solutions and measure outcomes

Package capability into solutions people use, and measure business results rather than model activity.

Goals

  • Define each solution around a business problem, an owner and a measure of success.
  • Establish a baseline before launch so improvement can be shown.
  • Track outcome, quality, cost and control metrics together.
  • Feed outcome evidence back into data, models and policy.

Decisions you will face

What is the unit of delivery?
A solution: a problem, the data, the agent, the workflow, the governance and the outcome. Use the same seven-part template every time so solutions are comparable.
Which outcome metric counts?
One primary business metric per solution, such as cycle time, cost per case or resolution quality, chosen with the business owner before launch.
How do you measure the baseline?
Measure the current process for a representative period before AI is introduced, or run a comparison group. Without a baseline you will argue about impact.
What cost view do you need?
Cost per outcome, not only cost per token: include infrastructure, review time, and the cost of errors caught late.
How do you scale a working solution?
Reuse the identity, policy and observability patterns, and change only the data, agent and workflow. Each new solution should be cheaper than the last.

Architecture notes

A solution is where capability becomes value. A consistent structure helps governance and learning: Problem, AI capability, Data, Agent, Workflow, Governance, Outcome. Writing every solution in that order forces the questions that matter, especially "what will we measure" and "what is the agent not allowed to do". The platform's solutions layer and industry pages apply the template.

Four families of metrics for every solution.
FamilyExamplesQuestion it answers
OutcomeCycle time, cost per case, resolution quality, revenue influencedDid the business result improve?
QualityAcceptance rate, edit rate, evaluation scores, escalation rateIs the AI doing the work well?
CostCost per outcome, GPU utilisation, review hoursIs it economical?
ControlPolicy denials, approvals, incidents, time to suspendIs it under control?

Do not stop at throughput. A solution that closes more tickets while quietly raising complaint rates has not improved the outcome. Pair every efficiency measure with a quality counter-measure. Report control metrics to the same audience as outcome metrics, so governance is seen as part of performance, not a tax on it.

Outcomes should flow back. Evidence from the solution (which answers were corrected, which cases escalated, which policy denials fired) improves retrieval, evaluation sets and policy. That is the closed intelligence loop: control, intelligence, agency, execution, outcomes, then evidence back into data and context. Read the closed intelligence loop deep dive for how the benefit accrues to the organisation, which gets smarter through its own AI use. For the ladder of expansion, see the platform hub and the AI estate inventory post.

Common pitfalls

  • Launching without a baseline, then relying on anecdote.
  • Counting usage (prompts, sessions) as value.
  • Reporting only the average. Look at the tail, where harm lives.
  • Treating each solution as a bespoke project with its own governance, which makes scaling slow and costly.
  • Never retiring solutions that no longer earn their keep.

Checklist

  • Solution written in the seven-part template.
  • Business owner and primary outcome metric agreed before build.
  • Baseline measured or a comparison group defined.
  • Metrics dashboard covers outcome, quality, cost and control.
  • Review cadence set with the business owner.
  • Feedback path from corrections and escalations into evaluations and retrieval.
  • Criteria for retiring or redesigning the solution.

Step 10 of 10

10Operate and learn

Run the platform as a living system: monitor it, review it, respond to incidents and let evidence improve it.

Goals

  • Operate AI with the same discipline as any production service: monitoring, incident response and change control.
  • Maintain a current inventory of models, agents, data flows and owners.
  • Review controls and requirements on a fixed cadence.
  • Turn operational evidence into better data, models, policies and solutions.

Decisions you will face

Who runs the platform day to day?
A platform team owns the runtime, policy engine and observability. Business owners own solutions. Security owns detection and response. Write the split down.
What is an AI incident?
Define it: a policy violation, a harmful or wrong action, a data exposure, a model outage with business impact. Each class needs a severity and a response path.
How often do you review?
Autonomy levels, permissions and approved models quarterly at minimum, and immediately after any incident or change of model or provider.
How do you handle change in the supply chain?
Track vendors, models and critical dependencies, watch for terms changes and deprecations, and rehearse the exit plan from step 1 on a schedule.
How do you keep up with regulation?
Assign someone to track it. Dates move: the EU AI Act high-risk deadlines changed in 2026. Record each obligation with its source and review date.

Architecture notes

Operations is where sovereignty is kept or lost. A platform that was well governed on launch day drifts: new agents appear, permissions accrete, models are swapped, data flows multiply. The countermeasure is an inventory and a rhythm. Our post on inventorying the AI estate describes how to find every model, agent and data flow, including the unsanctioned ones.

Monitoring and incident response

Monitor behaviour, not only uptime: policy denials by agent, approval rates, escalation volumes, anomalous tool calls, cost per workflow, evaluation scores on sampled traffic. Alerts should map to a runbook with an owner. The first response to an AI incident is usually containment: suspend the agent, revoke the credential, drop its autonomy level. Because you built a kill switch in step 5 and traceability in step 7, you can then establish what the agent did and which data it touched. The enterprise governance and risk post and AI compliance monitoring workflow are useful companions.

A review calendar

  • Weekly. Review alerts, denials and escalations; sample outputs for quality.
  • Monthly. Review cost, utilisation, evaluation trends and new agent requests.
  • Quarterly. Review autonomy levels, permissions, approved models, supplier register and the exit plan.
  • Annually, or on major change. Revisit the requirements from step 1 and the regulatory register.

Learning

Operational evidence is the raw material for improvement. Corrections made by reviewers become evaluation cases. Frequent escalations reveal missing context in the data layer or an over-tight policy. Cases where a model failed point to routing changes. This is the loop that makes the organisation smarter through its own AI use, and because the data, memory and evaluations live in your boundary, the benefit stays with you. The what is sovereign intelligence article puts this in the wider context, and the architecture reference shows how the parts connect.

Common pitfalls

  • Treating launch as the finish line, with no named owner for operations.
  • No inventory, so shadow agents and forgotten credentials accumulate.
  • Reviews that look at dashboards but never change a permission or a level.
  • Waiting for an incident to test the kill switch or the exit plan.
  • Letting a model or provider change land without re-running evaluations.

Checklist

  • Platform, business and security responsibilities written and assigned.
  • Inventory of models, agents, tools and data flows kept current.
  • AI incident classes, severities and runbooks defined.
  • Review calendar scheduled and attended.
  • Kill switch and exit plan rehearsed at least once a year.
  • Regulatory register with sources and review dates.
  • Operational feedback wired into evaluation sets, retrieval and policy.

Where to go next

The ten steps are a sequence, but they are also a loop. What you learn in operations changes your requirements, your data, your models and your policies, and the next pass is faster because the controls already exist. Capability without control is not enterprise-ready, and control without intelligence is not valuable. The aim is both.

Where to go next

When you are ready to build, start with Swfte through the entry point that fits your team, or talk to our team about a dedicated deployment.

Frequently asked questions

How long does it take to build a Sovereign Intelligence Platform?

It depends on scope, and we do not publish a universal figure. Most organisations start with one entry point, such as a knowledge assistant or a single governed agent, and expand layer by layer as they prove value. The land-and-expand path runs from an API through models, enterprise data, agents and workflows to dedicated infrastructure.

Do I need on-premises hardware to be sovereign?

No. Sovereignty is the ability to retain meaningful control over your AI estate, which is about control, not just where a server sits. Some workloads need on-premises or air-gapped deployment; many are well served by dedicated or private cloud. Decide per workload in step 1 and step 2.

In which order should the ten steps be done?

In the order given for a first pass, but governance decisions begin in step 1, not step 7. Teams often run steps 3 and 4 in parallel, and step 7 should be designed before any agent touches real systems in step 5.

Can I follow this guide with a mix of vendors and open-source tools?

Yes. The guide is vendor-neutral where it describes decisions, and the Swfte sections explain how the platform approaches each step. The tests that matter are portability, policy enforcement, auditability and the ability to exit.

Does following this guide make my organisation compliant?

No guide or platform does that on its own. The steps help you put in place technical controls, governance mechanisms and evidence that support your applicable regulatory, security and policy requirements. The exact posture depends on your use case, jurisdiction, deployment and configuration, and you should take legal advice for your situation.

Build it one step at a time

Start with one entry point. Add intelligence, agents, workflows and infrastructure as you prove value. Or read the step-by-step build guide and take the readiness assessment.

Ready to build with Swfte?

One platform for the agents, models and workflows your team ships. Free to start, no card required.