Deploy · Advanced

How to build an air-gapped AI environment

  • Time: Allow one to two days for the first environment, mostly waiting on downloads, transfers and approvals. Later updates take a few hours.
  • Cost: The tools are free. Your costs are the server hardware, storage for the staging area, and the time of whoever approves and moves each import.
  • Level: Advanced
On this page
  1. Short answer
  2. Before you start
  3. What you give up in an air gap
  4. 1. Define the boundary and the transfer rules
  5. 2. Download the model weights on the staging machine
  6. 3. Stage the container images and Python packages
  7. 4. Record checksums and pack the transfer set
  8. 5. Verify and load everything on the isolated side
  9. 6. Run the model with every online lookup switched off
  10. 7. Prove that nothing leaves the network
  11. 8. Plan updates, logging and rollback
  12. The same two-phase pattern elsewhere
  13. Troubleshooting
  14. Verify it worked
  15. Next steps
  16. FAQ
  17. How Swfte can help
  18. Sources and last verified

Short answer

An air-gapped AI environment is built in two halves. On a connected staging machine you download model weights, container images and Python wheels, record checksums and pack them. On the isolated side you verify the checksums, load everything, run the model with online lookups disabled, and test that no connection leaves the network. Updates repeat the same cycle, deliberately and on a schedule.

The steps at a glance

  1. Define the boundary and the transfer rules
  2. Download the model weights on the staging machine
  3. Stage the container images and Python packages
  4. Record checksums and pack the transfer set
  5. Verify and load everything on the isolated side
  6. Run the model with every online lookup switched off
  7. Prove that nothing leaves the network
  8. Plan updates, logging and rollback

Before you start

Who this is for

  • Platform and security engineers who must run a language model where prompts, documents and outputs cannot reach the internet.
  • Teams in regulated, defence-adjacent or industrial settings who already run isolated networks and now need to add an inference server to one.
  • Anyone who wants a repeatable way to bring open-weight models into a closed network and show an auditor how they got there.

Probably not for you if

Prerequisites

  • A connected staging machine with Docker, Python and the Hugging Face CLI, and enough free disk for the weights, the container images and a wheelhouse of Python packages. Budget about two bytes per parameter for a 16-bit model.
  • An isolated target server with the same CPU architecture and operating system family as the staging machine, an NVIDIA driver already installed, and Docker (or Ollama) installed from media you have approved.
  • An approved transfer route: removable media, a one-way data diode, or a controlled file drop. Write down who may use it and how it is logged.
  • A model and licence decision already made. Read the licence on the exact checkpoint before you carry it across, because you cannot look it up later from inside the gap.
  • A change-control record: who approves each import, and where the checksums are stored.
Time
Allow one to two days for the first environment, mostly waiting on downloads, transfers and approvals. Later updates take a few hours.
Cost
The tools are free. Your costs are the server hardware, storage for the staging area, and the time of whoever approves and moves each import.
Hardware
A GPU server sized for the model you choose (see [how to self-host an LLM](/how-to-self-host-an-llm) for the memory arithmetic), plus a staging machine with a fast disk.
Skill
Comfortable with Linux administration, Docker and networking basics.

Estimates are ours, not measurements, and move with your hardware, data and network.

What you give up in an air gap

An air gap buys isolation and costs convenience. You cannot pull a new model on a whim, call a hosted tool, fetch a package you forgot, or search the web from an agent. Anything an application used to fetch live must now be staged in advance or removed.

Be honest with your users about this before the build. A model that cannot browse or call external tools is still very useful for summarising, drafting, extraction and question answering over documents you place inside. It is not a replacement for a connected assistant, and teams who promise that tend to disappoint.

  1. Step 1Define the boundary and the transfer rules

    You end up with: A one-page statement of what is inside the gap, what may cross it, in which direction, and who signs.

    Write the boundary down before you download anything. Name the networks that count as inside, the single route by which files may enter, and whether anything may ever leave. Most teams allow files in and nothing out except reviewed reports. If you allow nothing out at all, say so, because it changes how you collect logs and error reports.

    List the four kinds of thing that will cross: model weights and tokenizer files, container images, Python packages, and operating system updates. Each needs an owner, a checksum record and an approval step. Treat them as a supply chain. A weight file is code you will run on a machine that holds sensitive data, and the file formats differ in risk. the Hugging Face security documentation warns that loading a pickle file can run arbitrary code, and that pickle is the default format for PyTorch weights. Prefer safetensors or GGUF, and load pickled checkpoints only from sources you trust.

    Decide how you will tell the people inside that something has changed. In an isolated network nobody reads release notes by accident, so an import ticket should say what was added, why, and which test it passed. That ticket is also the beginning of your evidence pack.

  2. Step 2Download the model weights on the staging machine

    You end up with: A directory holding the exact model files, pinned to a revision, with nothing else mixed in.

    Install the Hugging Face command line on the connected machine. The current documentation recommends a standalone installer, and pip install -U "huggingface_hub" also works. The command is hf, with hf download for fetching.

    Use --local-dir to put the files in a folder you control instead of the shared cache. The documentation notes that a .cache/huggingface/ folder is created inside it to track what was downloaded. For a GGUF model you usually want one quantisation, not the whole repository, so name the file or use --include to filter. Pin a revision with --revision so you can say exactly what you imported. Run with --dry-run first to see what would be fetched.

    Replace the placeholders below with the repository and file you chose. The repository name is yours to pick. Check its licence page before downloading, and record the commit hash that hf download resolves to.

    Install the CLI (Linux or macOS) · bash
    curl -LsSf https://hf.co/cli/install.sh | bash
    Preview, then download one GGUF file to a staging folder · bash
    hf download <org>/<repo> <file>.gguf --local-dir ./staging/model --dry-run
    hf download <org>/<repo> <file>.gguf --local-dir ./staging/model
    Or download a safetensors repository for vLLM · bash
    hf download <org>/<repo> --local-dir ./staging/model-safetensors

    Checked against: Hugging Face Hub CLI guide, huggingface_hub cache guide, Hugging Face Hub: pickle scanning and security

  3. Step 3Stage the container images and Python packages

    You end up with: Tar archives of every image you need and a folder of Python wheels that installs without an index.

    Pull the images you will run, then save them to tar archives. docker save writes one or several images to a file, and you can pipe it through gzip to shrink it. Pin images by an exact tag or digest, not latest. The vLLM documentation shows vllm/vllm-openai:latest in its example, which is fine for a first try on a connected machine and wrong for a record you will be asked about later.

    For Python packages, pip download fetches wheels without installing them. Give it a requirements file and a destination folder with -d. On the isolated side you will install with the package index switched off. The pip documentation says --no-index makes pip ignore the index and look only at --find-links locations, which is exactly the behaviour you want.

    Build the wheelhouse on a machine that matches the target: same operating system family, CPU architecture and Python version. pip does have --platform, --python-version and --only-binary options for downloading for a different target, but a matching staging machine avoids a long list of surprises.

    Save an image archive (replace the image with the exact tag you pinned) · bash
    docker pull <image>:<tag>
    docker save <image>:<tag> | gzip > ./staging/images/<name>.tar.gz
    Download every package in a requirements file into a wheelhouse · bash
    python -m pip download -r requirements.txt -d ./staging/wheelhouse

    Checked against: Docker image save and image load reference, pip download reference, vLLM Docker deployment

  4. Step 4Record checksums and pack the transfer set

    You end up with: One archive or folder with a checksum manifest, plus a copy of the manifest stored separately.

    Create a SHA-256 manifest of every file in the staging folder with sha256sum, and write the output to a file. On the other side you will check it with sha256sum -c. A manifest only proves the files did not change in transit. It does not prove they were safe to begin with, so keep the manifest, the repository name, the resolved revision and the licence text together in the import ticket.

    Store a second copy of the manifest by a different route from the media itself, for example on the approval ticket or on paper. If an attacker can alter the files on the media they can alter a manifest sitting next to them.

    Keep the transfer set small and single-purpose. One import per ticket makes it clear what changed when something misbehaves, and makes rollback a matter of deleting a folder.

    Write the manifest from inside the staging folder · bash
    cd staging
    find . -type f ! -name SHA256SUMS -print0 | xargs -0 sha256sum > SHA256SUMS

    Checked against: GNU coreutils: sha2 utilities

  5. Step 5Verify and load everything on the isolated side

    You end up with: Verified files in place, images loaded into Docker and Python packages installed from the wheelhouse.

    Copy the transfer set to the isolated server and run the checksum check first. Do not unpack, load or install anything that fails it. sha256sum -c prints a line per file and a final warning if any file differs, so a clean run is silent apart from the OK lines.

    Load the image archives with docker load. The Docker documentation shows both docker load < file.tar.gz and docker load --input file.tar, and it prints Loaded image: with the name and tag when it succeeds. Then install packages from the wheelhouse with the index off.

    If you will use Ollama, import the GGUF with a Modelfile. The Ollama documentation says a single GGUF file needs a Modelfile with a FROM line pointing at the file, then ollama create and ollama run. It also says Ollama does not quantise a GGUF during import, so the quantisation you downloaded is the one you get.

    Verify (stop here if any line says FAILED) · bash
    cd /import/staging
    sha256sum -c SHA256SUMS
    Load images and install packages without an index · bash
    docker load --input /import/staging/images/<name>.tar.gz
    python -m pip install --no-index --find-links=/import/staging/wheelhouse -r requirements.txt
    Import a GGUF into Ollama · text
    # Modelfile
    FROM /models/<file>.gguf
    Then build and test the model · bash
    ollama create my-model -f Modelfile
    ollama run my-model

    Checked against: GNU coreutils: sha2 utilities, Docker image save and image load reference, pip download reference, Ollama: importing models

  6. Step 6Run the model with every online lookup switched off

    You end up with: A working inference endpoint that loads from local files only and reports nothing home.

    Libraries built around the Hugging Face Hub try to reach it by default, even when the files are local. Set HF_HUB_OFFLINE=1. The documentation says that with it set, no HTTP calls are made to the Hub and only cached files are accessed, and that if a file is not in the cache an error is raised. That error is useful: it tells you what you forgot to stage.

    For vLLM, the troubleshooting documentation recommends downloading the model first with the hf CLI and passing the local path to vLLM. Mount the model directory into the container and point --model at the mounted path. vLLM also collects usage statistics by default. Its documentation gives two ways to turn that off: VLLM_NO_USAGE_STATS=1 or DO_NOT_TRACK=1, or creating a do_not_track file under ~/.config/vllm. In an isolated network the call would fail anyway, but a deliberate setting is better than a failure someone has to explain.

    Bind the server to an internal address only, and require an API key as you would on a connected network. How to self-host an LLM covers the memory sizing, authentication and TLS in detail.

    Serve a local safetensors model with vLLM, offline · bash
    docker run --runtime nvidia --gpus all \
      -v /models/model-safetensors:/models/model-safetensors:ro \
      --env "HF_HUB_OFFLINE=1" \
      --env "VLLM_NO_USAGE_STATS=1" \
      --env "DO_NOT_TRACK=1" \
      -p 127.0.0.1:8000:8000 \
      --ipc=host \
      <image>:<tag> \
      --model /models/model-safetensors

    Checked against: huggingface_hub environment variables, vLLM troubleshooting, vLLM usage stats collection, vLLM Docker deployment

  7. Step 7Prove that nothing leaves the network

    You end up with: Evidence, from the host and from the network edge, that the server makes no outbound connections.

    Believing the network is closed is not the same as showing it. Run two checks and keep the output. First, on the host, list established TCP connections with the processes that own them. ss -t state established lists established TCP sockets, -n stops it resolving names, and -p adds the process. Run it while the model is serving requests and during a restart, and look for anything that is not your own clients.

    Second, check at the edge. The firewall or switch between the server and the rest of the network should deny outbound traffic by default and log denied attempts. A log with zero denied outbound attempts from the model server across a week of use is stronger evidence than a one-off look at a terminal. If you do see attempts, they tell you which component is still trying to reach the internet, and you can fix the setting rather than rely on the firewall.

    For an extra isolation layer on a single container that does not need to serve other machines, Docker can start it with --network none, which creates only a loopback device. That suits batch jobs and test runs. A server that other machines must reach needs a network, so there you rely on the host and edge checks.

    List established TCP connections with owning processes, no name lookups · bash
    ss -tnp state established
    Run a throwaway container with only a loopback interface · bash
    docker run --rm --network none alpine:latest ip link show

    The second command should show only the loopback device

    1: lo: <LOOPBACK,UP,LOWER_UP> ...

    Checked against: ss(8) manual page, Docker none network driver

  8. Step 8Plan updates, logging and rollback

    You end up with: A written cycle for bringing in new models and patches, and a way to roll back.

    Decide how often you will import. A quiet air-gapped environment quietly falls behind: models, container images, drivers and operating system patches all age. Pick a cadence, such as monthly for patches and per request for new models, and make someone responsible for running it.

    Repeat steps 2 to 6 for every update, with a new ticket and a new manifest. Keep the previous model directory and image until the new one has passed your acceptance prompts, then remove it. Run a fixed set of test prompts before and after each import so you can see behaviour changes, not just that the process starts. How to validate your AI explains how to build that set.

    Log requests and responses inside the gap according to your data rules, and decide how logs are retained and who can read them. If nothing may leave, review logs inside the gap. If a reviewed summary may leave, agree its format in advance.

    What to keep for each import
    ItemWhy
    Repository, file name and resolved revisionLets you say exactly which weights are running
    Licence text and date readLicences change between releases and you cannot check later
    SHA256SUMS and the second copyShows the files arrived unchanged
    Test-prompt results before and afterShows what changed in behaviour
    Approver and dateAuditor evidence

The same two-phase pattern elsewhere

The two-phase shape in this guide is the common pattern. NVIDIA describes the same idea for its NIM containers in its air-gap deployment documentation: prepare assets on a machine with internet access, then run on the isolated machine with no outbound access and no API keys. If you use a vendor runtime, read its own air-gap page, because the specifics (which files, which environment variables) differ.

Troubleshooting

What you seeLikely causeFix
The server starts but fails with a message that a file cannot be found in the cache or that it cannot reach huggingface.coA tokenizer, config or generation file was not staged, or the server is still resolving a Hub name instead of a local path.Point --model at the mounted directory path, not at an org/name identifier, and check the directory holds config.json, the tokenizer files and every weight shard.
sha256sum -c prints FAILED for one or more filesThe file changed in transit, the copy was incomplete, or the manifest was written against a different path.Do not use the file. Copy it again from the staging machine, and make sure you check from the same relative directory the manifest was created in.
With HF_HUB_OFFLINE=1 an error says a file is missing from the cacheThis is the setting working as documented: only cached files are accessed.Go back to the staging machine, download the missing file with the same revision, add it to the manifest and repeat the import.
pip install --no-index reports that no matching distribution was foundThe wheelhouse was built on a machine with a different Python version, operating system or CPU architecture, or a dependency was not in the requirements file.Rebuild the wheelhouse on a machine that matches the target, using pip download -r requirements.txt -d wheelhouse, and test the install on a disconnected machine before the transfer.
docker load finishes but docker run says the image does not existYou saved or loaded under a different tag than you are running.Run docker image ls and use the exact repository and tag printed. Pin the tag when you save.
Ollama create succeeds but the model answers poorly or the quantisation is not what you expectedOllama does not quantise a GGUF on import, so the file you supplied is what runs.Download the quantisation you want from the start, or quantise on the staging machine with the tool that produced the GGUF, then re-import.
The edge firewall logs denied outbound attempts from the model serverSome component still tries to reach the internet: telemetry, a package manager, a model hub lookup.Match the destination to a process with ss -tnp, then turn off the feature or set the offline variable. Keep the firewall deny rule either way.

Verify it worked

Next steps

Related guides

Frequently asked questions

What is an air-gapped AI environment?

It is a network with no path to the public internet where a model runs on local hardware, so prompts, documents and outputs never cross a security boundary. Everything the model needs, from weights to software packages, is carried in by an approved route and checked on arrival.

Can you run an LLM offline without internet?

Yes. Download the weights and software once on a connected machine, move them to the offline machine, and run them with tools such as Ollama or vLLM. Set HF_HUB_OFFLINE=1 so libraries stop trying to reach the Hugging Face Hub and fail clearly if a file is missing.

How do I get model weights into an air-gapped network?

Download them on a connected staging machine with hf download, record SHA-256 checksums, copy them over an approved route such as removable media, and verify the checksums on the other side before loading anything. Record the licence and revision with the import.

How do I prove an AI server has no internet access?

Use two checks. On the host, run ss -tnp state established to list connections and their processes. At the network edge, deny outbound traffic by default and log denied attempts. Keep both outputs as evidence over a period of real use, not just a single look.

How do I update models in an air-gapped network?

Repeat the staging process on a schedule: download, checksum, transfer, verify, load, test. Keep the previous version until the new one passes your test prompts, so you can roll back. Use one import ticket per change so the record stays clear.

Does vLLM send data when it runs offline?

vLLM collects usage statistics by default. Its documentation lists what is collected and gives opt-outs: set VLLM_NO_USAGE_STATS=1 or DO_NOT_TRACK=1, or create a do_not_track file in ~/.config/vllm. Set one deliberately rather than relying on the missing network to block it.

How Swfte can help

You can build everything above with open-source tools and no vendor. If you would rather not run the infrastructure yourself, Swfte offers dedicated and self-deployed options, and the gateway can sit in front of a model you host.

Fully air-gapped installation of Swfte products is designed for and available on request: <air-gapped install availability - founder to fill>. Nothing in this guide depends on it.

Missing a step or found a command that no longer works? Tell us, or request a how-to.

Sources and last verified

Commands, versions and facts in this guide were checked against the sources below on . Tools change quickly: if something differs from what you see, trust the official documentation and let us know.

  1. Hugging Face Hub CLI guide: installer and pip install commands, hf download, --local-dir, --include, --dry-run, offline note
  2. huggingface_hub environment variables: HF_HUB_OFFLINE behaviour and accepted boolean values, HF_HOME and HF_HUB_CACHE
  3. huggingface_hub cache guide: cache layout (blobs, refs, snapshots), hf cache verify, incomplete snapshot errors offline
  4. vLLM Docker deployment: docker run command for vllm/vllm-openai with --runtime nvidia, --gpus all, --ipc=host, --model
  5. vLLM troubleshooting: recommendation to download the model with the hf CLI and pass the local path (page dated 21 September 2026)
  6. vLLM usage stats collection: VLLM_NO_USAGE_STATS, DO_NOT_TRACK and the do_not_track file opt-outs
  7. Ollama: importing models: Modelfile FROM line for GGUF, ollama create, ollama run, no quantisation on import
  8. Docker image save and image load reference: docker save to a file or through gzip; docker load with < and --input (from the load reference page)
  9. Docker none network driver: docker run --network none creates only a loopback device
  10. pip download reference: pip download -r, -d, --no-index, --find-links, platform options
  11. GNU coreutils: sha2 utilities: sha256sum and the -c / --check verification option
  12. ss(8) manual page: options -t, -n, -p and the state established filter
  13. Hugging Face Hub: pickle scanning and security: loading a pickle file can execute arbitrary code; pickle is the default PyTorch weight format; safetensors as an alternative
  14. NVIDIA NIM air-gap deployment: the two-phase prepare-then-run pattern (read via search result summary, not a full page fetch)

Topics

  • air-gapped
  • offline
  • model supply chain
  • vLLM
  • Ollama
  • egress testing

Machine-readable copies: this guide as markdown, index of all guides (JSON). Canonical address: https://www.swfte.com/how-to-build-an-air-gapped-ai-environment.

Ready to build with Swfte?

One platform for the agents, models and workflows your team ships. Free to start, no card required.