Sovereignty · Data

AI Data Sovereignty: Location, Access, Processing, Transfer and Retention

Updated 2026-10-06 · 6 min read

Short answer:Data sovereignty for AI means the organisation decides where its data lives, who can access it, how it is processed, whether it moves and how long it is kept. Residency, meaning storage in a chosen region, is only the first of those five. Prompts, retrieved context, embeddings and logs count as data too.

On this page

What are the five controls of data sovereignty?

Data sovereignty is the first of seven kinds of control in a Sovereign Intelligence estate, described at AI sovereignty. It breaks into five controls. You decide where data lives, who can reach it, how it is processed, whether it moves, and how long it stays.

ControlThe questionExample decision
LocationWhere is data stored and processed?Region, country, on a device, in a dedicated environment
AccessWho and what can read or write it?User roles, agent identities, support staff of the provider
ProcessingWhat is done to it, and by which model or service?Which classes of data may reach a hosted model API
TransferDoes it move, and under which safeguards?Cross-border flows, sub-processors, model-provider calls
RetentionHow long does it exist, in every copy?Deletion of prompts, embeddings and logs on a schedule

Location gets most of the attention because it is easy to state and easy to sell. The other four are where control is more often lost.

Why is data residency not the same as data sovereignty?

Residency says where bytes are stored. Sovereignty asks who ultimately controls them. The two diverge whenever the operator of the infrastructure is subject to a legal system different from the one the data belongs to.

The US CLOUD Act of 2018 is the usual example. It allows US authorities to compel US-based providers to disclose data in their possession, custody or control regardless of where it is stored. The test follows the provider, not the data centre. A CSIS analysis (opens in a new tab) lays out the mechanism and the transatlantic tension, including the comity process a provider can use to challenge an order that conflicts with foreign law. That tension is with the GDPR, whose Article 48 says a foreign court order is not, by itself, a lawful ground for transferring EU personal data.

The practical consequence is that a region setting is a start, not an answer. The questions that matter are: who operates the environment, what law can reach them, who holds the encryption keys and what happens to the data if a request arrives. Infrastructure-level versions of these questions are in infrastructure sovereignty.

What counts as data in an AI system?

Traditional data programmes list databases and files. AI adds several data types that are easy to miss, and each needs the same five controls.

  • Prompts and conversations. They contain whatever a person typed, including customer details and unreleased plans.
  • Retrieved context. Passages pulled from documents at answer time. Their sensitivity equals that of the source, but they are copied into a model call.
  • Embeddings and vector indexes. Numeric representations of your content. They are derived from sensitive data and can leak information about it, so classify them with their source. See data and context.
  • Memory. What an agent chooses to remember between sessions.
  • Model outputs. Drafts, summaries and decisions, which may themselves be personal or confidential.
  • Logs and traces. The record of what happened. Ironically these often hold more sensitive content than the original system, and they are frequently retained longest.
  • Fine-tuning and evaluation sets. Curated extracts from your data. Once a model is tailored on them, the data's influence outlives the dataset.

A common gap is a data classification policy that covers databases and ignores the vector store, because nobody thought of it as a copy. Add AI artefacts to the register explicitly.

How do classification and routing enforce it?

The workable pattern is to keep three or four classes of data and, for each, state which places may process it. The policy is the decision. The platform's job is to enforce it.

Illustrative routing policy. The classes and rules are examples, not a recommendation for your organisation.
ClassMay be processed onTypical controls
PublicAny approved model, including hosted APIsStandard logging
InternalApproved hosted models under a data-handling agreementPrompt and output logging, retention limit
ConfidentialModels you host in a controlled environment, or on-deviceRestricted access, filtered outputs, no provider retention
RestrictedDedicated or on-premises environment onlyHuman approval for exports, strict key control

Our write-up of six enforcement layers for data sovereignty is a good model of what enforce means in practice, and the model gateway view shows how requests can be routed by policy to different models. For the model-selection side of the same decision, see model sovereignty.

What role do encryption keys play?

Who holds the keys often decides who really controls the data. If the provider holds them, the provider can be compelled or can err. If you hold them, the provider's copy is unreadable without your cooperation. This is why the Commission's framework treats control of cryptographic access as part of its data and AI objective.

Be precise about what a platform offers. For Swfte, the trust centre states the position: customer data is stored in AWS eu-west-1 (Ireland) today, with encryption at rest on the databases and storage in use. Customer-managed encryption keys are designed for on dedicated deployments and are not generally available today. Other data locations are agreed per dedicated deployment, scoped with you rather than offered self-serve. On-device mode processes prompts with a local model on a Mac and does not send them to Swfte or model providers. Read that page rather than a summary of it, since it also lists what is still in progress.

How should retention work for AI artefacts?

Retention is the control teams most often leave at the vendor's default. Decide it deliberately, per artefact type, and write it into the contract and the system's Trust Profile (which has a retention field; see Trust Profile).

  1. Set separate periods for prompts, outputs, embeddings, memory and logs. They do not need the same period.
  2. Propagate deletion. Removing a source document should remove its chunks and embeddings, and invalidate cached answers built from it.
  3. Reconcile with audit needs. Records kept for oversight can conflict with minimisation. Resolve this by recording references and decisions rather than full content wherever the audit purpose allows.
  4. Check provider retention. Ask each model provider what it keeps and for how long. Zero-retention terms are a contract matter, and Swfte lists its own agreements with providers as in progress on the trust centre.
  5. Test deletion. Ask for evidence that a deleted item is gone from indexes and backups within the stated window.

How do EU rules interact with this?

Three EU instruments shape the conversation, and none of them is a certificate you can hold. The GDPR governs personal data and cross-border transfers. The Data Act, applicable from 12 September 2025, gives users rights over data generated by connected products and, in its Chapter VI, obliges cloud providers to remove switching barriers; see infrastructure sovereignty for that part. The AI Act adds record-keeping and oversight duties for high-risk systems, which depend on being able to keep and read logs; see governance sovereignty.

Swfte's approach is compliance-by-design. It provides the technical controls, governance mechanisms and evidence required to deploy AI within an organisation's applicable regulatory, security and policy requirements. The exact posture depends on the customer's use case, jurisdiction, deployment and configuration. This is general information, not legal advice. Take advice for your own situation, and see the EU AI Act overview for the timeline.

A data sovereignty checklist for AI

  • Every AI data type (prompts, context, embeddings, memory, outputs, logs, training sets) is on the data register with a class and an owner.
  • Each class has a written list of places allowed to process it, and the platform enforces the list.
  • You know who operates each environment and which legal systems can reach them.
  • You know who holds encryption keys for each store, and what the provider can read.
  • Every sub-processor and model provider in the data path is listed. See the Swfte sub-processors page for the draft list.
  • Retention is set per artefact type, propagates on deletion and is tested.
  • You can export your data and indexes in a readable format.
  • Access by agents is scoped to the identity and the task, not to a shared credential.

Work through these in order as step 1 and step 3 of the build guide, and ask the chief data officer questions of any vendor you assess.

Frequently asked questions

What is the difference between data residency and data sovereignty?

Residency is where data is stored. Sovereignty is who controls it: access, processing, transfer, retention and the legal reach over the operator. Data can be resident in your country and still be exposed to a foreign legal demand if the operator is subject to it.

Do prompts and embeddings count as personal or sensitive data?

They can. Prompts contain whatever people type. Embeddings derive from your content and can leak information about it. Classify both with their source data and apply the same location, access, transfer and retention controls.

Does keeping data in the EU solve the CLOUD Act issue?

Not by itself. The Act follows the provider's possession, custody or control, not the storage location. Whether it applies depends on the provider's legal ties and structure. Assess this per provider with legal counsel, along with who holds the encryption keys.

Where does Swfte store customer data today?

In AWS eu-west-1 (Ireland), per the trust centre. Other locations and customer-managed keys are designed for dedicated deployments, scoped per engagement, and are not generally available today. Check the trust centre for current status.

Can I keep sensitive prompts off third-party model providers?

That is the purpose of classification and routing: restricted classes go only to models you host or to on-device models. Swfte's on-device mode processes prompts with a local model and does not send them to Swfte or model providers.

Sources cited

Put data sovereignty into practice

Start with one entry point. Add intelligence, agents, workflows and infrastructure as you prove value. Or read the step-by-step build guide and take the readiness assessment.

Ready to build with Swfte?

One platform for the agents, models and workflows your team ships. Free to start, no card required.