← The journal
Security

Enclosed AI: Offline AI for Regulated Teams

Enclosed AI keeps models, data and agents inside a boundary you control: three levels of enclosure.

Swfte Journal / Security

Enclosed AI is AI that runs entirely inside a boundary you control, so prompts, documents, retrieved context and outputs do not cross it. Nothing about the phrase is standardized. Searches for it return pages about air-gapped AI, offline AI and sovereign AI, which is a useful clue: people use the word for a property rather than a product. The property is that the model comes to the data, instead of the data traveling to the model.

This post defines three levels of enclosure, lists what has to be inside the boundary for the claim to hold, and explains how regulated teams, such as legal, healthcare, finance and public sector, decide how tight a boundary they need.

What does "enclosed" actually mean?

An AI system is enclosed when you can state, for every request, that the data stayed within a defined perimeter: a device, a network segment or a facility. That statement needs three things to be true.

  1. The model runs inside the perimeter. You cannot enclose an API. If the model is a hosted service, the prompt travels to it by definition. Enclosure only works with models you can run yourself, which means open-weight models or models licensed for self-hosting.
  2. Everything the model needs is also inside. Document parsing, embeddings, the vector index, the user interface and the authentication system. One external call in that chain, even for something minor like a font or a license check, breaks the claim.
  3. Egress is controlled and verifiable. You can show that the system does not call out, and that attempts to do so are blocked or logged.

What are the three levels of enclosure?

LevelBoundaryWhat it protects againstWhat it costs you
On-deviceThe user's laptop or workstationContent leaving the endpoint at allLimited to models the hardware can run; no shared index
Private network, controlled egressYour network or a dedicated environment, with deny-by-default outbound rulesThird-party processing of content; most exfiltration pathsYou operate or contract for the infrastructure; some allowlisted updates
Air-gappedA network with no connection to the outsideRemote access and network exfiltrationManual updates, physical transfer, slower iteration

Most regulated teams do not need the third level for most work. Matching the level to the data class is the actual skill. A legal team's privileged matter files may justify on-device or private-network AI. A defense, critical-infrastructure or classified environment may require the air gap. The air-gapped LLM deployment checklist covers the tightest level in detail.

What has to be inside the boundary?

Use this as an audit list. If any item is outside, the system is not enclosed.

  • The language model weights and the inference server.
  • The embedding model and the vector store.
  • Document ingestion: OCR, parsing and chunking.
  • The chat or agent interface and its static assets.
  • Identity and access management for the users of the system.
  • Logs, metrics and the audit trail, with no outbound telemetry.
  • The package mirrors, container registry and model registry used for updates.
  • The tools and connectors agents can call, restricted to internal systems.

One detail that trips people up: inference engines and model libraries often have telemetry, update checks or online fetches that are on by default. The practical guides for air-gapped vLLM, for example, list specific environment variables to turn off Hugging Face lookups and usage statistics, and they stress that these are necessary but not sufficient. Verify by running the system with outbound traffic blocked. If it fails, you found an egress path.

Which regulated teams use enclosed AI, and why?

Vendor pages in this space name defense contractors, critical infrastructure operators, healthcare providers, finance, legal and public-sector bodies as typical users. The reasons differ.

  • Legal and professional services. Privileged and confidential client material. The risk is confidentiality duties and client contract terms rather than a statute.
  • Healthcare. Patient data subject to data protection rules, with a high cost of a breach. Whether a given use is allowed depends on your jurisdiction and your agreements.
  • Financial services. Customer data, trading information and the supervisory expectation that you manage third-party risk. For European banks, the EU stack post covers DORA's exit-plan expectations.
  • Public sector and defense. Classified or sensitive data, procurement rules and in some cases a legal ban on external processing.
  • Industrial and energy. Operational technology networks that are segmented from the internet on purpose.

None of this means a regulation requires air-gapped AI. Most regulations are outcome-based, and enclosure is one control that makes the outcome easier to prove.

Is offline AI good enough?

For many tasks, yes. Open-weight models have closed much of the gap with hosted frontier models for drafting, summarizing, extraction and question answering over your own documents. Small models run on a modern laptop, and mid-sized ones run on a single server GPU. Our posts on running Qwen locally, Gemma 4 sizing and open weights versus proprietary models cover where the line sits now.

What you give up is the very top of the capability curve and the convenience of automatic updates. For tasks that need a frontier model, the choice is to route those requests, with the data class permitting, through a governed gateway, or to accept the limit.

What goes wrong inside the boundary?

Isolation removes external exposure, not internal risk.

  • Permission leaks. An enclosed assistant that indexes everything with an administrator account lets any user ask about anything. Retrieval must respect source permissions.
  • Prompt injection. The OWASP Top 10 for LLM Applications 2025 lists prompt injection first and sensitive information disclosure second. A closed network does not stop a malicious document from steering an agent.
  • Unlogged actions. An enclosed agent that can write to internal systems needs approval steps and a record. The AI DMZ architecture post describes controlled ingress, auditable execution and egress control.
  • Stale models. An air-gapped model drifts out of date. Plan a signed, versioned update process.

What does a practical enclosed setup look like?

A common progression:

  1. Start on-device for individuals. Content stays on the endpoint and there is no infrastructure to run.
  2. Add a private-network inference server for shared team workloads, with an internal gateway so applications call one endpoint.
  3. Keep a short, governed list of exceptions that may reach hosted models, with data-class rules enforced at the gateway.
  4. Move the most sensitive environments to an air-gapped deployment with a defined transfer procedure.

Swfte's Cortex is the first step. It is a governed AI desktop that runs on the laptop by default, its on-device model answers with no internet connection, and it warns when a request would leave the device. Today it ships for Macs with Apple Silicon, with other platforms to follow. The second and fourth steps are the domain of dedicated cloud, which includes isolated and air-gapped options, and the gateway is BuildX. Studio and Nexus cover building agents and watching what they do. The infrastructure and sovereignty pages explain how these fit.

Swfte is designed to provide technical controls, governance mechanisms and evidence for deploying AI within your own regulatory, security and policy requirements. The exact posture depends on your use case, jurisdiction, deployment and configuration. To discuss an enclosed deployment, contact the team. Related: sovereign AI versus private cloud and confidential computing for LLMs.

Frequently asked questions

What is enclosed AI?

It is AI that runs inside a boundary you control, a device, a network or a facility, so that prompts, documents and outputs do not leave it. It requires a self-hosted model and control of every supporting component.

Is enclosed AI the same as air-gapped AI?

Air-gapped AI is the strictest form, with no network connection to the outside. Enclosed AI is broader and includes on-device and private-network designs with controlled egress.

Can offline AI work with our own documents?

Yes. Retrieval over internal documents runs locally with a local embedding model and vector store, as long as every component of the pipeline is inside the boundary.

Does a regulation require enclosed AI?

Rarely by name. Most rules set outcomes for confidentiality, residency and third-party risk. Enclosure is a control that can help demonstrate the outcome. Check with counsel for your sector.

How do we keep an offline model up to date?

Treat updates as releases: download on a connected machine, verify hashes, carry across approved media, run your offline evaluations, then promote. See the air-gapped checklist.

Related: Swfte Connect is the model gateway, designed to run in your own cloud or data centre; see self-deploying Connect.

Further reading: Custom Domain Models vs Frontier APIs for Regulated Teams.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Ready to build with Swfte?

One platform for the agents, models and workflows your team ships. Free to start, no card required.