# How to build an air-gapped AI environment

Canonical: https://www.swfte.com/how-to-build-an-air-gapped-ai-environment
Last verified: 2026-10-06
Difficulty: Advanced
Time: Allow one to two days for the first environment, mostly waiting on downloads, transfers and approvals. Later updates take a few hours.
Cost: The tools are free. Your costs are the server hardware, storage for the staging area, and the time of whoever approves and moves each import.
Hardware: A GPU server sized for the model you choose (see [how to self-host an LLM](/how-to-self-host-an-llm) for the memory arithmetic), plus a staging machine with a fast disk.

## Short answer

An air-gapped AI environment is built in two halves. On a connected staging machine you download model weights, container images and Python wheels, record checksums and pack them. On the isolated side you verify the checksums, load everything, run the model with online lookups disabled, and test that no connection leaves the network. Updates repeat the same cycle, deliberately and on a schedule.

## Who this is for

- Platform and security engineers who must run a language model where prompts, documents and outputs cannot reach the internet.
- Teams in regulated, defence-adjacent or industrial settings who already run isolated networks and now need to add an inference server to one.
- Anyone who wants a repeatable way to bring open-weight models into a closed network and show an auditor how they got there.

Not for:
- People who only want data to stay in the EU or in one region. That is a residency question, not an isolation one: see [how to deploy an LLM in the EU](https://www.swfte.com/how-to-deploy-an-llm-in-the-eu).
- Anyone who just wants a model on a laptop. [How to run LLMs locally](https://www.swfte.com/how-to-run-llms-locally) is shorter and does not need a transfer process.

## Prerequisites

- A connected staging machine with Docker, Python and the Hugging Face CLI, and enough free disk for the weights, the container images and a wheelhouse of Python packages. Budget about two bytes per parameter for a 16-bit model.
- An isolated target server with the same CPU architecture and operating system family as the staging machine, an NVIDIA driver already installed, and Docker (or Ollama) installed from media you have approved.
- An approved transfer route: removable media, a one-way data diode, or a controlled file drop. Write down who may use it and how it is logged.
- A model and licence decision already made. Read the licence on the exact checkpoint before you carry it across, because you cannot look it up later from inside the gap.
- A change-control record: who approves each import, and where the checksums are stored.

## What you give up in an air gap

An air gap buys isolation and costs convenience. You cannot pull a new model on a whim, call a hosted tool, fetch a package you forgot, or search the web from an agent. Anything an application used to fetch live must now be staged in advance or removed.

Be honest with your users about this before the build. A model that cannot browse or call external tools is still very useful for summarising, drafting, extraction and question answering over documents you place inside. It is not a replacement for a connected assistant, and teams who promise that tend to disappoint.

## Steps

### Step 1: Define the boundary and the transfer rules

Outcome: A one-page statement of what is inside the gap, what may cross it, in which direction, and who signs.

Write the boundary down before you download anything. Name the networks that count as inside, the single route by which files may enter, and whether anything may ever leave. Most teams allow files in and nothing out except reviewed reports. If you allow nothing out at all, say so, because it changes how you collect logs and error reports.

List the four kinds of thing that will cross: model weights and tokenizer files, container images, Python packages, and operating system updates. Each needs an owner, a checksum record and an approval step. Treat them as a supply chain. A weight file is code you will run on a machine that holds sensitive data, and the file formats differ in risk. the Hugging Face security documentation warns that loading a pickle file can run arbitrary code, and that pickle is the default format for PyTorch weights. Prefer safetensors or GGUF, and load pickled checkpoints only from sources you trust.

Decide how you will tell the people inside that something has changed. In an isolated network nobody reads release notes by accident, so an import ticket should say what was added, why, and which test it passed. That ticket is also the beginning of your evidence pack.

> NOTE: This is not legal or accreditation advice. If your network is classified or subject to a national scheme, its own rules for media handling and transfer override anything here.

### Step 2: Download the model weights on the staging machine

Outcome: A directory holding the exact model files, pinned to a revision, with nothing else mixed in.

Install the Hugging Face command line on the connected machine. The current documentation recommends a standalone installer, and `pip install -U "huggingface_hub"` also works. The command is `hf`, with `hf download` for fetching.

Use `--local-dir` to put the files in a folder you control instead of the shared cache. The documentation notes that a `.cache/huggingface/` folder is created inside it to track what was downloaded. For a GGUF model you usually want one quantisation, not the whole repository, so name the file or use `--include` to filter. Pin a revision with `--revision` so you can say exactly what you imported. Run with `--dry-run` first to see what would be fetched.

Replace the placeholders below with the repository and file you chose. The repository name is yours to pick. Check its licence page before downloading, and record the commit hash that `hf download` resolves to.

Install the CLI (Linux or macOS):

```bash
curl -LsSf https://hf.co/cli/install.sh | bash
```

Preview, then download one GGUF file to a staging folder:

```bash
hf download <org>/<repo> <file>.gguf --local-dir ./staging/model --dry-run
hf download <org>/<repo> <file>.gguf --local-dir ./staging/model
```

Or download a safetensors repository for vLLM:

```bash
hf download <org>/<repo> --local-dir ./staging/model-safetensors
```

> TIP: Download the tokenizer and config files too. A model directory without `tokenizer.json` or `config.json` loads fine on the staging machine, where missing files are fetched silently, and fails offline.

### Step 3: Stage the container images and Python packages

Outcome: Tar archives of every image you need and a folder of Python wheels that installs without an index.

Pull the images you will run, then save them to tar archives. `docker save` writes one or several images to a file, and you can pipe it through `gzip` to shrink it. Pin images by an exact tag or digest, not `latest`. The vLLM documentation shows `vllm/vllm-openai:latest` in its example, which is fine for a first try on a connected machine and wrong for a record you will be asked about later.

For Python packages, `pip download` fetches wheels without installing them. Give it a requirements file and a destination folder with `-d`. On the isolated side you will install with the package index switched off. The pip documentation says `--no-index` makes pip ignore the index and look only at `--find-links` locations, which is exactly the behaviour you want.

Build the wheelhouse on a machine that matches the target: same operating system family, CPU architecture and Python version. pip does have `--platform`, `--python-version` and `--only-binary` options for downloading for a different target, but a matching staging machine avoids a long list of surprises.

Save an image archive (replace the image with the exact tag you pinned):

```bash
docker pull <image>:<tag>
docker save <image>:<tag> | gzip > ./staging/images/<name>.tar.gz
```

Download every package in a requirements file into a wheelhouse:

```bash
python -m pip download -r requirements.txt -d ./staging/wheelhouse
```

### Step 4: Record checksums and pack the transfer set

Outcome: One archive or folder with a checksum manifest, plus a copy of the manifest stored separately.

Create a SHA-256 manifest of every file in the staging folder with `sha256sum`, and write the output to a file. On the other side you will check it with `sha256sum -c`. A manifest only proves the files did not change in transit. It does not prove they were safe to begin with, so keep the manifest, the repository name, the resolved revision and the licence text together in the import ticket.

Store a second copy of the manifest by a different route from the media itself, for example on the approval ticket or on paper. If an attacker can alter the files on the media they can alter a manifest sitting next to them.

Keep the transfer set small and single-purpose. One import per ticket makes it clear what changed when something misbehaves, and makes rollback a matter of deleting a folder.

Write the manifest from inside the staging folder:

```bash
cd staging
find . -type f ! -name SHA256SUMS -print0 | xargs -0 sha256sum > SHA256SUMS
```

### Step 5: Verify and load everything on the isolated side

Outcome: Verified files in place, images loaded into Docker and Python packages installed from the wheelhouse.

Copy the transfer set to the isolated server and run the checksum check first. Do not unpack, load or install anything that fails it. `sha256sum -c` prints a line per file and a final warning if any file differs, so a clean run is silent apart from the OK lines.

Load the image archives with `docker load`. The Docker documentation shows both `docker load < file.tar.gz` and `docker load --input file.tar`, and it prints `Loaded image:` with the name and tag when it succeeds. Then install packages from the wheelhouse with the index off.

If you will use Ollama, import the GGUF with a Modelfile. The Ollama documentation says a single GGUF file needs a Modelfile with a `FROM` line pointing at the file, then `ollama create` and `ollama run`. It also says Ollama does not quantise a GGUF during import, so the quantisation you downloaded is the one you get.

Verify (stop here if any line says FAILED):

```bash
cd /import/staging
sha256sum -c SHA256SUMS
```

Load images and install packages without an index:

```bash
docker load --input /import/staging/images/<name>.tar.gz
python -m pip install --no-index --find-links=/import/staging/wheelhouse -r requirements.txt
```

Import a GGUF into Ollama:

```text
# Modelfile
FROM /models/<file>.gguf
```

Then build and test the model:

```bash
ollama create my-model -f Modelfile
ollama run my-model
```

### Step 6: Run the model with every online lookup switched off

Outcome: A working inference endpoint that loads from local files only and reports nothing home.

Libraries built around the Hugging Face Hub try to reach it by default, even when the files are local. Set `HF_HUB_OFFLINE=1`. The documentation says that with it set, no HTTP calls are made to the Hub and only cached files are accessed, and that if a file is not in the cache an error is raised. That error is useful: it tells you what you forgot to stage.

For vLLM, the troubleshooting documentation recommends downloading the model first with the `hf` CLI and passing the local path to vLLM. Mount the model directory into the container and point `--model` at the mounted path. vLLM also collects usage statistics by default. Its documentation gives two ways to turn that off: `VLLM_NO_USAGE_STATS=1` or `DO_NOT_TRACK=1`, or creating a `do_not_track` file under `~/.config/vllm`. In an isolated network the call would fail anyway, but a deliberate setting is better than a failure someone has to explain.

Bind the server to an internal address only, and require an API key as you would on a connected network. [How to self-host an LLM](https://www.swfte.com/how-to-self-host-an-llm) covers the memory sizing, authentication and TLS in detail.

Serve a local safetensors model with vLLM, offline:

```bash
docker run --runtime nvidia --gpus all \
  -v /models/model-safetensors:/models/model-safetensors:ro \
  --env "HF_HUB_OFFLINE=1" \
  --env "VLLM_NO_USAGE_STATS=1" \
  --env "DO_NOT_TRACK=1" \
  -p 127.0.0.1:8000:8000 \
  --ipc=host \
  <image>:<tag> \
  --model /models/model-safetensors
```

> WARNING: The vLLM documentation example publishes `-p 8000:8000`, which listens on every interface. The command above binds to 127.0.0.1 so a reverse proxy decides who can reach it.

### Step 7: Prove that nothing leaves the network

Outcome: Evidence, from the host and from the network edge, that the server makes no outbound connections.

Believing the network is closed is not the same as showing it. Run two checks and keep the output. First, on the host, list established TCP connections with the processes that own them. `ss -t state established` lists established TCP sockets, `-n` stops it resolving names, and `-p` adds the process. Run it while the model is serving requests and during a restart, and look for anything that is not your own clients.

Second, check at the edge. The firewall or switch between the server and the rest of the network should deny outbound traffic by default and log denied attempts. A log with zero denied outbound attempts from the model server across a week of use is stronger evidence than a one-off look at a terminal. If you do see attempts, they tell you which component is still trying to reach the internet, and you can fix the setting rather than rely on the firewall.

For an extra isolation layer on a single container that does not need to serve other machines, Docker can start it with `--network none`, which creates only a loopback device. That suits batch jobs and test runs. A server that other machines must reach needs a network, so there you rely on the host and edge checks.

List established TCP connections with owning processes, no name lookups:

```bash
ss -tnp state established
```

Run a throwaway container with only a loopback interface:

```bash
docker run --rm --network none alpine:latest ip link show
```

The second command should show only the loopback device:

```text
1: lo: <LOOPBACK,UP,LOWER_UP> ...
```

### Step 8: Plan updates, logging and rollback

Outcome: A written cycle for bringing in new models and patches, and a way to roll back.

Decide how often you will import. A quiet air-gapped environment quietly falls behind: models, container images, drivers and operating system patches all age. Pick a cadence, such as monthly for patches and per request for new models, and make someone responsible for running it.

Repeat steps 2 to 6 for every update, with a new ticket and a new manifest. Keep the previous model directory and image until the new one has passed your acceptance prompts, then remove it. Run a fixed set of test prompts before and after each import so you can see behaviour changes, not just that the process starts. [How to validate your AI](https://www.swfte.com/how-to-validate-your-ai) explains how to build that set.

Log requests and responses inside the gap according to your data rules, and decide how logs are retained and who can read them. If nothing may leave, review logs inside the gap. If a reviewed summary may leave, agree its format in advance.

**What to keep for each import**

| Item | Why |
| --- | --- |
| Repository, file name and resolved revision | Lets you say exactly which weights are running |
| Licence text and date read | Licences change between releases and you cannot check later |
| SHA256SUMS and the second copy | Shows the files arrived unchanged |
| Test-prompt results before and after | Shows what changed in behaviour |
| Approver and date | Auditor evidence |

## The same two-phase pattern elsewhere

The two-phase shape in this guide is the common pattern. NVIDIA describes the same idea for its NIM containers in its air-gap deployment documentation: prepare assets on a machine with internet access, then run on the isolated machine with no outbound access and no API keys. If you use a vendor runtime, read its own air-gap page, because the specifics (which files, which environment variables) differ.

## Troubleshooting

| Symptom | Likely cause | Fix |
| --- | --- | --- |
| The server starts but fails with a message that a file cannot be found in the cache or that it cannot reach huggingface.co | A tokenizer, config or generation file was not staged, or the server is still resolving a Hub name instead of a local path. | Point `--model` at the mounted directory path, not at an `org/name` identifier, and check the directory holds `config.json`, the tokenizer files and every weight shard. |
| `sha256sum -c` prints FAILED for one or more files | The file changed in transit, the copy was incomplete, or the manifest was written against a different path. | Do not use the file. Copy it again from the staging machine, and make sure you check from the same relative directory the manifest was created in. |
| With HF_HUB_OFFLINE=1 an error says a file is missing from the cache | This is the setting working as documented: only cached files are accessed. | Go back to the staging machine, download the missing file with the same revision, add it to the manifest and repeat the import. |
| `pip install --no-index` reports that no matching distribution was found | The wheelhouse was built on a machine with a different Python version, operating system or CPU architecture, or a dependency was not in the requirements file. | Rebuild the wheelhouse on a machine that matches the target, using `pip download -r requirements.txt -d wheelhouse`, and test the install on a disconnected machine before the transfer. |
| `docker load` finishes but `docker run` says the image does not exist | You saved or loaded under a different tag than you are running. | Run `docker image ls` and use the exact repository and tag printed. Pin the tag when you save. |
| Ollama create succeeds but the model answers poorly or the quantisation is not what you expected | Ollama does not quantise a GGUF on import, so the file you supplied is what runs. | Download the quantisation you want from the start, or quantise on the staging machine with the tool that produced the GGUF, then re-import. |
| The edge firewall logs denied outbound attempts from the model server | Some component still tries to reach the internet: telemetry, a package manager, a model hub lookup. | Match the destination to a process with `ss -tnp`, then turn off the feature or set the offline variable. Keep the firewall deny rule either way. |

## Verify it worked

- [ ] Every file in the transfer set passed `sha256sum -c` on the isolated side, and a second copy of the manifest matches.
- [ ] The model answers a fixed set of test prompts with `HF_HUB_OFFLINE=1` set and no route to the internet.
- [ ] `ss -tnp state established` run during use shows only connections to and from machines you expect.
- [ ] The edge firewall shows no permitted outbound connection from the model server, and its log of denied attempts is empty or explained.
- [ ] Usage statistics are switched off for vLLM and any other component that has such a setting.
- [ ] The import ticket records the repository, revision, licence, checksums, test results and approver.
- [ ] You have restored the previous version once, to prove rollback works.

## Next steps

- [How to self-host an LLM](https://www.swfte.com/how-to-self-host-an-llm): size the GPU, add an API key and TLS, and set up health checks for the server inside the gap
- [How to validate your AI](https://www.swfte.com/how-to-validate-your-ai): build the before-and-after test set you need for each import
- [How to evaluate an open-source LLM](https://www.swfte.com/how-to-evaluate-an-open-source-llm): check licence and task fit before you carry a model across
- [Air-gapped LLM deployment checklist](https://www.swfte.com/blog/air-gapped-llm-deployment-practical-checklist): a shorter checklist version for planning conversations

## FAQ

### What is an air-gapped AI environment?

It is a network with no path to the public internet where a model runs on local hardware, so prompts, documents and outputs never cross a security boundary. Everything the model needs, from weights to software packages, is carried in by an approved route and checked on arrival.

### Can you run an LLM offline without internet?

Yes. Download the weights and software once on a connected machine, move them to the offline machine, and run them with tools such as Ollama or vLLM. Set `HF_HUB_OFFLINE=1` so libraries stop trying to reach the Hugging Face Hub and fail clearly if a file is missing.

### How do I get model weights into an air-gapped network?

Download them on a connected staging machine with `hf download`, record SHA-256 checksums, copy them over an approved route such as removable media, and verify the checksums on the other side before loading anything. Record the licence and revision with the import.

### How do I prove an AI server has no internet access?

Use two checks. On the host, run `ss -tnp state established` to list connections and their processes. At the network edge, deny outbound traffic by default and log denied attempts. Keep both outputs as evidence over a period of real use, not just a single look.

### How do I update models in an air-gapped network?

Repeat the staging process on a schedule: download, checksum, transfer, verify, load, test. Keep the previous version until the new one passes your test prompts, so you can roll back. Use one import ticket per change so the record stays clear.

### Does vLLM send data when it runs offline?

vLLM collects usage statistics by default. Its documentation lists what is collected and gives opt-outs: set `VLLM_NO_USAGE_STATS=1` or `DO_NOT_TRACK=1`, or create a `do_not_track` file in `~/.config/vllm`. Set one deliberately rather than relying on the missing network to block it.

## How Swfte can help

You can build everything above with open-source tools and no vendor. If you would rather not run the infrastructure yourself, Swfte offers dedicated and self-deployed options, and the gateway can sit in front of a model you host.

- [Dedicated cloud](https://www.swfte.com/dedicated-cloud): infrastructure set aside for your organisation
- [Platform infrastructure layer](https://www.swfte.com/platform/infrastructure): how Swfte describes control over where and how AI runs
- [Connect self-deploy](https://www.swfte.com/products/connect/self-deploy): how the gateway can be deployed in your own environment

Fully air-gapped installation of Swfte products is designed for and available on request: <air-gapped install availability - founder to fill>. Nothing in this guide depends on it.

## Sources

- [Hugging Face Hub CLI guide](https://huggingface.co/docs/huggingface_hub/guides/cli): installer and pip install commands, hf download, --local-dir, --include, --dry-run, offline note
- [huggingface_hub environment variables](https://huggingface.co/docs/huggingface_hub/package_reference/environment_variables): HF_HUB_OFFLINE behaviour and accepted boolean values, HF_HOME and HF_HUB_CACHE
- [huggingface_hub cache guide](https://huggingface.co/docs/huggingface_hub/guides/manage-cache): cache layout (blobs, refs, snapshots), hf cache verify, incomplete snapshot errors offline
- [vLLM Docker deployment](https://docs.vllm.ai/en/stable/deployment/docker.html): docker run command for vllm/vllm-openai with --runtime nvidia, --gpus all, --ipc=host, --model
- [vLLM troubleshooting](https://docs.vllm.ai/en/stable/usage/troubleshooting.html): recommendation to download the model with the hf CLI and pass the local path (page dated 21 September 2026)
- [vLLM usage stats collection](https://docs.vllm.ai/en/stable/usage/usage_stats.html): VLLM_NO_USAGE_STATS, DO_NOT_TRACK and the do_not_track file opt-outs
- [Ollama: importing models](https://docs.ollama.com/import): Modelfile FROM line for GGUF, ollama create, ollama run, no quantisation on import
- [Docker image save and image load reference](https://docs.docker.com/reference/cli/docker/image/save/): docker save to a file or through gzip; docker load with < and --input (from the load reference page)
- [Docker none network driver](https://docs.docker.com/engine/network/drivers/none/): docker run --network none creates only a loopback device
- [pip download reference](https://pip.pypa.io/en/stable/cli/pip_download/): pip download -r, -d, --no-index, --find-links, platform options
- [GNU coreutils: sha2 utilities](https://www.gnu.org/software/coreutils/manual/html_node/sha2-utilities.html): sha256sum and the -c / --check verification option
- [ss(8) manual page](https://man7.org/linux/man-pages/man8/ss.8.html): options -t, -n, -p and the state established filter
- [Hugging Face Hub: pickle scanning and security](https://huggingface.co/docs/hub/security-pickle): loading a pickle file can execute arbitrary code; pickle is the default PyTorch weight format; safetensors as an alternative
- [NVIDIA NIM air-gap deployment](https://docs.nvidia.com/nim/large-language-models/2.0.2/deployment/air-gap-deployment.html): the two-phase prepare-then-run pattern (read via search result summary, not a full page fetch)

Last verified against these sources on 2026-10-06.
