← The journal
Engineering

Air-Gapped LLM Deployment: A Practical Checklist

An air-gapped LLM checklist: stage weights, verify them, switch off hidden egress and test that it runs offline.

Swfte Journal / Engineering

An air-gapped LLM deployment works if you can pull the network cable and every user-facing function still runs. Getting there is a staging problem, a hidden-egress problem and a sizing problem. You stage everything on a connected machine, verify it, carry it across the gap, switch off every default that phones home, and then prove the isolation by testing with outbound traffic blocked.

This is a working checklist for platform engineers. It uses vLLM as the serving example because it is open source and common, but the structure applies to any inference server. For the wider concept and the lighter levels of enclosure, start with enclosed AI for regulated teams.

What does an air-gapped LLM need that a normal one does not?

A normal LLM deployment pulls things from the internet at install time and often at run time: container images, Python packages, model weights, tokenizer files, license checks and telemetry. Inside an air gap, none of that is available, so each dependency has to be moved in by hand, once, deliberately. That is the whole difference, and it is also why teams underestimate the work. The model is the easy part.

Phase 1: What do you stage on the connected side?

Do the downloading on a machine that has internet access, and treat it as a build step.

  • Container images. Save the inference server image and every sidecar. Reference images by version tag or digest from your internal registry, never by a floating tag such as latest, which fails at run time when it tries to resolve.
  • Python wheels and system packages. Download dependencies for the exact Python and CUDA versions on the target.
  • Model weights. Download the full model directory: weights, tokenizer files and configuration. Serve a local path rather than a hub repository name.
  • Embedding and reranking models, if you use retrieval. These are easy to forget.
  • Any custom model code. Some architectures need code that the inference server fetches when you allow remote code. Pre-stage it, and review it.
  • A manifest. Record every file with a SHA-256 hash, the source, the version, the license and the date.

Large models are large. A 70-billion-parameter model in 16-bit precision needs roughly 140 GB for the weights alone, since each parameter takes two bytes. Plan removable storage or an approved transfer path accordingly. OpenAI's gpt-oss-120b, released under Apache 2.0, ships natively quantized so the checkpoint is around 61 GB, which is one reason such releases suit constrained environments. Check the license of every model you plan to carry in.

Phase 2: How do you cross the gap safely?

  • Bundle and sign. One archive containing the images, wheels, weights and manifest, signed with a key you control. Practitioner guides recommend a customer-controlled trust root.
  • Verify on arrival. Check the signature and every hash against the manifest, and where the model author publishes hashes, against theirs.
  • Scan. Run your normal malware and software-composition checks on the imported content.
  • Import to a registry, not a shared folder. Load images into an internal container registry and weights into an internal model registry. Inference workers pull from these, never from the public internet.
  • Record provenance. For each model version, store the source, hash, import date, importing engineer and the evaluation results that qualified it.

Phase 3: How do you switch off hidden egress?

This is where most "air-gapped" deployments quietly are not. Defaults in common libraries try to reach the network. Typical switches for a vLLM and Hugging Face based stack, as described in practitioner guides to air-gapped vLLM, look like this.

export HF_HUB_OFFLINE=1
export TRANSFORMERS_OFFLINE=1
export HF_HUB_DISABLE_TELEMETRY=1
export HF_HUB_DISABLE_UPDATE_CHECK=1
export VLLM_NO_USAGE_STATS=1

Treat the list as a starting point, not a guarantee. HF_HUB_OFFLINE governs only the Hugging Face hub library, and other components have their own switches. Check the documentation for your exact versions, and watch for:

  • Interactive API documentation pages that load assets from a CDN.
  • Per-request features that fetch URLs, such as multimodal inputs.
  • Tokenizer lookups when files are not pre-staged.
  • Update checks in command-line tools.
  • Telemetry in the web interface, vector database and observability agents.
  • Authentication providers that call out for token validation.

The real control is the network, not the environment variables. Variables reduce noise. Firewall rules and the absence of a route enforce the property.

Phase 4: How do you size the hardware?

Start from the arithmetic, then confirm by measurement.

ItemRule of thumb
Weights memoryParameters multiplied by bytes per parameter: about 2 bytes at 16-bit, 1 at 8-bit, 0.5 at 4-bit
KV cacheGrows with context length and concurrent requests; often the real limit in production
HeadroomActivations, runtime overhead and fragmentation; leave a margin rather than filling the card
ConcurrencyTest with your own prompt lengths; throughput depends on batching

vLLM is Apache-2.0 licensed, exposes an OpenAI-compatible API server and supports continuous batching and multiple quantization formats on NVIDIA, AMD and other hardware, according to its repository. Our continuous batching deep dive explains the scheduling, and the serving frameworks comparison helps if you are choosing between engines. For GPU purchasing, see the GPU strategy post, and for the economics, the TCO analysis.

Phase 5: What else must be inside?

An LLM server on its own is not a workspace. Add, inside the boundary:

  • An internal gateway so applications call one endpoint and models can be swapped.
  • A vector store and an embedding model.
  • An identity provider that works offline, with single sign-on and role-based access.
  • Logging and metrics that stay inside, with an audit trail of who asked what and which sources were retrieved.
  • Internal PKI and encrypted traffic between components.
  • A policy layer for agents: what tools they may call and which actions need human approval. See the AI DMZ pattern.

How do you prove it is air-gapped?

Test, and keep the evidence.

  1. Start the whole stack with the network physically disconnected, or with all outbound traffic dropped at the firewall. It must come up and answer.
  2. Capture traffic on every host for a representative day of use. Any outbound attempt is a finding.
  3. Run your offline evaluation suite and compare with the baseline from the connected environment.
  4. Attempt the abuse cases: prompt injection through a document, a user asking about a file they should not see, an agent asked to call an unapproved tool.
  5. Record results with dates and versions.

How do you update an air-gapped model?

Treat an update as a release. Build and verify on the connected side, sign, carry across, verify again, stage in the registry under a new version label, run offline evaluations, then promote. Keep the previous version so you can roll back. Decide a cadence in advance, because the risk of never updating is a stale model and unpatched software.

Common mistakes

  • Treating the model as the whole project and discovering the other twenty dependencies at install time.
  • Pulling latest tags.
  • Believing one environment variable closes all egress.
  • Forgetting the embedding model and the reranker.
  • No update procedure.
  • No agent policy: an offline agent with unrestricted tools is still dangerous.
  • No audit trail, so the first incident cannot be investigated.

Where does Swfte fit?

Swfte's dedicated cloud includes isolated and air-gapped deployment options, BuildX provides the gateway layer across models, and Studio and Nexus build agents and enforce policy on what they do. Cortex covers the endpoint, with an on-device model that answers with no internet connection. The infrastructure and agents pages describe the layers.

Swfte is designed to provide technical controls, governance mechanisms and evidence for deploying AI within your own regulatory, security and policy requirements. The exact posture depends on your use case, jurisdiction, deployment and configuration. If you are planning an air-gapped build, talk to the team. You may also want the self-hosted LLM stack guide and confidential computing for LLMs.

Frequently asked questions

What is an air-gapped LLM?

It is a language model deployed on infrastructure with no connection to external networks, so prompts, documents and outputs cannot leave. It must run on a model you host, with all dependencies staged inside.

Can vLLM run fully offline?

Yes, with preparation. You pre-stage weights and tokenizer files, serve from a local path, use internal image references and switch off the library features that call out. Verify by running with outbound traffic blocked.

How much GPU memory do I need?

Weights need roughly parameters times bytes per parameter, then add KV cache and overhead. Quantization reduces weights memory. Measure with your own context lengths and concurrency.

Is Ollama suitable for an air-gapped environment?

For small deployments and workstations, it can be, since it runs models locally and exposes a local API. For multi-user, high-throughput serving, teams usually choose a dedicated inference server. See the stack guide.

Is air-gapping required by regulation?

Rarely by name. It is a control that makes confidentiality and residency outcomes easier to demonstrate. Check your sector's rules with counsel.

Related: Swfte Connect is the model gateway, designed to run in your own cloud or data centre; see self-deploying Connect.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.