Deploy an Open-Source LLM in the EU, Step by Step
Deploy an open-source LLM on EU infrastructure with vLLM: pin the model, serve it privately, gate and monitor.
Running your own model in the EU sounds like a large project and is mostly a sequence of small, checkable steps. This post walks through one concrete path: choose and pin a model, bring it onto EU infrastructure, serve it with vLLM, put it behind a gateway, test it, and operate it. The commands are illustrative; adapt names, sizes and flags to your model and hardware, and check the current documentation for the version you install.
If you want the principles behind each step, the deploy models hub and the guides on deploying an LLM in the EU and self-hosted inference go further. This post is the runbook.
Step 0: Decide what "EU" has to mean for you
"In the EU" is not one requirement. Before you pick a region, write down which of these you actually need, because they lead to different infrastructure:
- Residency. Prompts, outputs, logs and backups stored and processed in an EU region.
- Jurisdiction. The operator of the infrastructure is itself under EU law, with no non-EU parent that can be compelled to disclose. This is a legal question; get advice for your situation.
- Access. Who can administer the machines, and from where.
- Dependencies. No lab-side model API, telemetry endpoint or package registry that quietly sends data elsewhere.
Open weights help with the last one: the model is files and inference is a process on hardware you control, so there is no model-lab API in the loop. They do not settle the first three on their own.
Step 1: Choose the model and read its licence for EU use
Shortlist by task, languages, context length and hardware fit. Then read the licence and acceptable use policy that ship with the exact checkpoint:
- Qwen's open-weight releases, Gemma 4 and most recent Mistral releases are Apache 2.0, and recent DeepSeek releases such as R1 are MIT.
- Kimi K2 uses a modified MIT with an attribution requirement above 100 million monthly active users or 20 million US dollars in monthly revenue.
- Llama is under a custom licence, and the Llama 4 acceptable use policy states that rights for its multimodal models are not granted to EU-domiciled individuals or EU-headquartered companies. Meta's FAQ says end users of a product that incorporates such a model are not affected, and that a company headquartered outside the EU may distribute products containing it. Have counsel read the current text before an EU company builds on it.
This is orientation, not legal advice. Keep the licence file with the artifact.
Step 2: Size the hardware
Weights cost about two bytes per parameter in bf16, one in FP8 and half a byte at 4-bit. A 70-billion-parameter dense model is therefore about 140 GB, 70 GB or 35 GB of weights respectively, before KV cache and activations. Add KV cache for your context length and concurrency, then leave headroom. A mixture-of-experts model needs all experts in memory even though few are active per token, so size for total parameters.
Pick GPUs from the GPU reference by memory and bandwidth, and decide the quantization now, because it changes both fit and behavior. If you will quantize, the quantized artifact is a new candidate and goes through the gate in Step 5.
Step 3: Pin, download and verify the weights
Download by commit hash, not by branch name:
hf download <repo_id> --revision <full-40-char-commit-hash> --local-dir ./model
The Hugging Face documentation is explicit that the revision should be the full commit hash. Then:
# Prefer safetensors; investigate anything else
ls ./model | grep -E "\.(bin|pt|pth|pkl)$" || echo "no pickle-based files"
# Record a hash for every file and keep the manifest with the artifact
(cd ./model && find . -type f -print0 | sort -z | xargs -0 sha256sum) > model.sha256
Do not load models that require remote code. Flags such as trust_remote_code=True execute code from the repository, and safetensors protects only the weight file, not the code around it. Copy the artifact to storage inside your EU environment and verify model.sha256 there. If the publisher offers a signature, verify it; the OpenSSF Model Signing specification covers weights, configs and tokenizers as one unit. More on this in model supply-chain security.
Step 4: Serve it with vLLM, privately
vLLM is the common default for shared GPU serving. It provides an OpenAI-compatible server with continuous batching, paged KV memory, prefix caching, tensor parallelism and quantization support. Hugging Face has put Text Generation Inference in maintenance mode and recommends vLLM or SGLang for new endpoints, so start there. A starting command, adapted to a model that spans two GPUs:
vllm serve ./model \
--served-model-name my-model \
--tensor-parallel-size 2 \
--max-model-len 32768 \
--gpu-memory-utilization 0.90 \
--api-key "$VLLM_API_KEY" \
--host 0.0.0.0 --port 8000
A few things to get right:
- Cap the context length at what the application needs. Engines reserve memory against it, so serving the model's maximum when you need a tenth of it wastes capacity.
- The API key is not a perimeter. vLLM's documentation says the key authenticates the
/v1family of paths, so other endpoints (such as health and metrics) are not covered by it. Put the server on a private network with no public ingress and make the gateway its only client. - Pin the engine version and the container image digest. An engine upgrade changes behavior and goes through the gate like any other change.
- Use the built-in operations endpoints.
/healthis a health check and/metricsexposes Prometheus-compatible metrics, which is enough to build dashboards and autoscaling signals.
Smoke test from inside the network:
curl -s http://model-host:8000/v1/chat/completions \
-H "Authorization: Bearer $VLLM_API_KEY" -H "Content-Type: application/json" \
-d '{"model":"my-model","messages":[{"role":"user","content":"Reply with the single word: ready"}]}'
Step 5: Evaluate, then harden, before anyone uses it
A model that answers a smoke test has not been tested. Before promotion, run:
- Your own task suite, scored against the baseline you use today.
- A safety suite: harmful-request refusal, over-refusal of benign requests, jailbreak resistance and prompt-injection cases. The OWASP Top 10 for LLM Applications 2025 is a good taxonomy, with prompt injection first.
- Language suites in each EU language you serve, compared with English. Safety behavior often degrades outside English.
- A latency and cost run on the target hardware at realistic concurrency.
Hardening means runtime guardrails (input and output checks, tool-permission limits, human approval for consequential actions), a fixed system policy, and the approved-model list enforced at the gateway. The full method is in how to evaluate an open-source LLM before production and the red-teaming checklist. Our testing methodology describes the gate we apply.
Step 6: Put it behind a gateway
Do not point applications at the engine. Put a gateway in front so that:
- one API serves every workload, and the model behind a route can change without rebuilding agents;
- approved-model policy is enforced per workload and data classification;
- requests are logged, with redaction and in-region retention;
- cost is tracked per route;
- traffic can fail over to another approved model and shift gradually between revisions.
Swfte's Connect is that gateway. Routing is also your rollback lever: keep the previous approved revision deployed until the new one has run on live traffic, then moving traffic back is a configuration change.
Step 7: Keep logs, telemetry and backups in-region
The model can be fully local while the evidence of its use leaks out through side channels. Check, explicitly:
- the log store, the metrics store and the tracing backend are in the EU region;
- backups and snapshots, including of the weights volume, are in-region;
- no component sends usage telemetry to a vendor endpoint, and package and container registries are mirrored where the environment is restricted;
- redaction runs before prompts reach a log.
Write each decision down as policy rather than leaving it as a default.
Step 8: Monitor and keep the evidence
Watch time to first token, tokens per second, queue depth, KV-cache utilization, cost per completed task, refusal rate and guardrail hits. Re-run the evaluation suite on every change that alters behavior: a new quantization, engine upgrade, system prompt, adapter or base revision. Keep the record of what ran: model and commit hash, licence, manifest, engine version, image digest, policy set, evaluation runs and the person who approved the promotion.
That record is the practical meaning of audit-ready. It does not settle your position under the EU AI Act or GDPR by itself: the exact posture depends on your use case, jurisdiction, deployment and configuration, and you should take legal advice. What it does is give you the evidence to answer the questions you will be asked.
The short version
- Define what EU means for you: residency, jurisdiction, access, dependencies.
- Choose the model; read its licence and acceptable use policy for EU use.
- Size memory from parameters, precision, context and concurrency.
- Pin the commit, use safetensors, verify hashes, keep remote code off.
- Serve with vLLM on a private network; the gateway is the only client.
- Evaluate and harden before promotion, including EU languages.
- Route through a gateway with approved-model policy and rollback.
- Keep logs, telemetry and backups in-region; monitor and record.
For a managed version of this loop on dedicated infrastructure, see deploy models with Swfte and dedicated cloud. Related reading: data sovereignty and AI, the EU sovereign AI stack for banks and the public sector, and the vLLM continuous batching deep dive.