Deployment guide
Deploy an LLM on-premise: offline weights and air-gapped inference
How to run an LLM inside your own data center, including with no outbound network.
On-premises deployment is the strongest answer to the question where does the model run, and the most demanding one to operate. Hardware has to be procured, weights have to arrive without a live connection to a model hub, and every update becomes a logistics exercise. Done well it gives you a model that physically cannot send a prompt anywhere. This guide covers the parts that go wrong: getting weights in safely, proving nothing phones home, and keeping the system patched and evaluated over time.
The on-premise path
Size and procure
Work out memory from the model, precision and context length (see the self-hosted inference guide), then add redundancy. Check interconnect needs if the model spans several GPUs. Accelerator lead times are your plan’s risk, so order early.
Build the artifact outside
On a connected staging host, download the model at a pinned full commit hash, verify its checksums, scan the files, run the evaluation gate, and package the approved weights with the licence file and a manifest.
Move it across the boundary
Transfer the bundle through your approved path (a data diode, a scanned removable medium or a one-way file gateway), then re-verify the hashes on the inside. A signature or hash list that travels separately from the bundle is what lets the inside trust the outside.
Serve with no outbound path
Run the engine from the local copy with offline mode on, and block egress at the network layer as well. Do not rely on a setting alone.
Operate and update
Patch the OS, drivers and engine on a schedule, and re-run the evaluation gate on every model or engine change before promotion. Keep the previous approved bundle for rollback.
Getting weights in without trusting the internet
Download by commit, not by branch. Branches move; a full 40-character commit hash names exactly one set of files. With the Hugging Face tooling, download with the revision set to that hash, then set the offline environment variable (HF_HUB_OFFLINE=1) on the serving host so the client uses only the local cache and fails rather than reaching out. Pre-downloading in a build step and shipping the cache is the usual pattern.
Prefer safetensors files. Pickle-based formats can execute arbitrary code when loaded, which is exactly the wrong property for the inside of an air gap. Safetensors was independently audited by Trail of Bits, with no critical flaw leading to arbitrary code execution found, but it protects only the weight file: it does not stop a tampered tensor, and it does not cover loading flags such as trust_remote_code, which should stay off unless you have reviewed the code.
For provenance, record a hash manifest and, where the publisher supports it, verify a signature. Model signing under the OpenSSF Model Signing specification produces a detached signature over weights, configs and tokenizers as one unit. Hashes prove the files are the ones you tested; they do not prove the model is benign, which is what the evaluation gate is for.
Proving nothing phones home
- Run the serving host in a network segment with no default route; allow only the gateway and the monitoring collector.
- Switch the client libraries to offline mode and treat any attempted outbound request as an alert.
- Disable usage telemetry in every component (engine, UI tooling, package managers) and verify it by capturing egress attempts in a test.
- Keep remote code execution off for model loading, and pin all container images by digest.
- Log what ran: model hash, engine version, image digest, policy set. That record is also your audit trail.
Hybrid: governed on-prem with a managed edge
Few organizations run everything air-gapped. A common shape is a private model on-premises for sensitive workloads and the Connect gateway routing other workloads to approved external or managed models, with one policy and one audit trail across both. The same models, agents and workflows can move between environments because the policy travels with the workload rather than being rebuilt per site.
Swfte supports dedicated deployment from an isolated VPC to bare metal in your own data center. On-premises, hybrid and air-gapped patterns are part of the platform’s stated direction; what is available for your site is agreed per engagement. <on-premises and air-gapped deployment options — founder to fill>.
Frequently asked questions
Can an LLM run fully air-gapped?
Yes. Open weights are files and the engine is software, so inference needs no internet once the weights, engine and dependencies are inside. The difficulty is update logistics and verifying what you bring in.
How do I get model weights into an air-gapped network safely?
Download at a pinned full commit hash on a staging host, verify checksums, scan, evaluate, package with a manifest, transfer through your approved path and re-verify hashes inside. Prefer safetensors and verify signatures where available.
What is HF_HUB_OFFLINE?
An environment variable for the Hugging Face client. Set to 1, it uses only the local cache and does not make network calls.
Do I still need an evaluation gate on-premise?
Yes. Hash checks show the files are the ones you tested; they do not show the model is safe. Run capability, safety and multilingual suites before the bundle crosses the boundary.
Can I mix on-premise and managed models?
Yes. Put the Connect gateway in front so policy, routing and the audit trail are the same across both.