Short answer
To self-host an LLM, pick an open-weight model whose weights and KV cache fit in GPU memory, run it with vLLM in Docker bound to localhost, require an API key, and put a reverse proxy with TLS in front that exposes only the /v1 paths. Check /health and /metrics from inside the host. Pin the image tag and test every upgrade against your own prompts before switching.
The steps at a glance
- Work out the memory you need
- Prepare the host: driver, Docker and the NVIDIA runtime
- Choose a licence-clean model and download it in advance
- Start vLLM in Docker, bound to localhost, with an API key
- Test the API from the server itself
- Put a TLS reverse proxy in front and expose only /v1
- Add health checks and metrics
- Tune for your load, and add GPUs only when the numbers say so
- Plan upgrades and rollback before you need them
Before you start
Who this is for
- Engineers who need a private model endpoint for an application, an internal tool or a team, and who are comfortable with Linux, Docker and a terminal.
- Teams that have outgrown a laptop setup and want several people or services sharing one GPU server.
- Security-minded teams who want prompts to stay on infrastructure they control.
Probably not for you if
- People who only want a model on their own laptop. Start with how to run LLMs locally, which is simpler.
- Teams whose usage is a few thousand requests a month. A hosted API is almost always cheaper and needs no operations.
Prerequisites
- A Linux server with an NVIDIA GPU of compute capability 7.5 or higher (the vLLM documentation lists T4, RTX 20-series, A100, L4, H100 and B200 as examples) and root or sudo access.
- The NVIDIA driver installed from your distribution, plus Docker. The NVIDIA documentation says the driver must be in place before you configure the container toolkit.
- A domain name you control for the public or internal address, or your own certificate authority if the server has no public DNS.
- Enough free local disk for the weights (roughly two bytes per parameter for a 16-bit model) and about 20 GB spare for the container image.
- Basic comfort reading logs and editing a text config file.
- Time
- About 2 to 3 hours for a first working endpoint; add a day for hardening, monitoring and a load test.
- Cost
- The software is free (vLLM is Apache-2.0). You pay for the GPU server, power or rental, and engineering time to run it.
- Hardware
- One NVIDIA GPU with 24 GB of memory runs an 8-billion-parameter model in 16-bit with a modest context. Larger models need more memory or several GPUs.
- Skill
- Comfortable with Linux, Docker and HTTP.
Estimates are ours, not measurements, and move with your hardware, data and network.
Which engine: vLLM, Ollama or TGI?
This guide uses vLLM because it is built for shared serving: it exposes an OpenAI-compatible server and is Apache-2.0 licensed. Ollama is the simpler tool for one person or a small group, and our own stack-layers post draws the same line. Ollama binds to 127.0.0.1 on port 11434 by default, and its FAQ describes OLLAMA_NUM_PARALLEL (default 1) and OLLAMA_MAX_LOADED_MODELS for tuning concurrency.
Text Generation Inference (TGI) needs a specific note. On 6 October 2026 the Hugging Face README says TGI "is now in maintenance mode" and will accept only minor bug fixes, documentation improvements and lightweight maintenance. It recommends vLLM, SGLang, and local engines such as llama.cpp or MLX going forward. If you already run TGI, nothing breaks today, but a new deployment should start elsewhere. Our vLLM versus TGI comparison predates that notice, so read its TGI column with that in mind.
| Engine | Best fit | Licence | Notes |
|---|---|---|---|
| vLLM | Shared GPU server, many concurrent users | Apache-2.0 | OpenAI-compatible server; Linux; NVIDIA compute capability 7.5+ |
| Ollama | Laptop, workstation, small team | MIT | Default bind 127.0.0.1:11434; simple setup; lower concurrency focus |
| TGI | Existing deployments only | Apache-2.0 | Maintenance mode per its README; new work is pointed to vLLM or SGLang |
Step 1Work out the memory you need
You end up with: A memory budget that tells you which GPU, and how much context and concurrency, you can afford.
Before you rent or buy anything, do the arithmetic. A model needs memory for its weights and, on top of that, a KV cache that grows with every token of every request in flight. Most "it crashed on startup" reports are a budget that was never written down.
Weights cost about two bytes per parameter in 16-bit, one byte in 8-bit and half a byte in 4-bit. The KV cache costs, per token, 2 (keys and values) x layers x key-value heads x head dimension x bytes per value. Read those numbers from the model's
config.json. For Qwen3-8B the file lists 36 layers, 8 key-value heads and a head dimension of 128, in bfloat16. That gives 2 x 36 x 8 x 128 x 2 = 147,456 bytes, about 144 KiB per token.Now apply it. The model card gives 8.2 billion parameters, so the 16-bit weights are about 16.4 GB. One sequence at 32,768 tokens needs 147,456 x 32,768 bytes, about 4.5 GiB of cache. On a 24 GB card, vLLM's default
--gpu-memory-utilizationof 0.92 leaves about 22 GB, so roughly 5.7 GB is left after the weights. That is an upper bound of about 38,000 tokens in flight across all users, before activations and overhead. Eight users at 4,000 tokens each fits; one user at 40,000 tokens does not.Weights-only memory for an 8.2 billion parameter model (Qwen3-8B), by precision Precision Bytes per parameter Weights Fits on a 24 GB card with cache room? 16-bit (bf16) 2 about 16.4 GB Yes, with about 5.7 GB for cache at 0.92 utilisation 8-bit 1 about 8.2 GB Yes, with much more cache room 4-bit 0.5 about 4.1 GB Yes, but check quality before you rely on it Checked against: Qwen3-8B config.json, Qwen3-8B model card, vLLM engine arguments
Step 2Prepare the host: driver, Docker and the NVIDIA runtime
You end up with: A container can see the GPU, proven by nvidia-smi running inside Docker.
vLLM runs on Linux. Install the NVIDIA driver with your distribution's package manager first, then Docker, then the NVIDIA Container Toolkit. The toolkit is what lets a container use the GPU.
Once the toolkit is installed, register it with Docker and restart the daemon. Then run the sample workload from the NVIDIA documentation. If it prints the usual GPU table, the host is ready. If it fails, fix that now; every later error will look like a vLLM problem when it is a driver or runtime problem.
Register the NVIDIA runtime with Docker and restart it · bash sudo nvidia-ctk runtime configure --runtime=docker sudo systemctl restart dockerProve a container can see the GPU · bash docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smiChecked against: NVIDIA Container Toolkit install guide, vLLM GPU installation
Step 3Choose a licence-clean model and download it in advance
You end up with: The weights sit on local disk, and you know exactly which revision they are.
Pick a model whose licence you have read. Qwen3-8B is labelled apache-2.0 on its Hugging Face card, which allows commercial use with notice and attribution. If you are choosing between several candidates, do that properly first: see how to evaluate an open-source LLM.
Download the weights before you start the server. The vLLM troubleshooting page says that a hang during download is fixed by downloading with the
hfcommand first and passing the local path to vLLM, and that loading from a network filesystem can hang, so use a local disk. Write down the commit hash shown on the model's "Files and versions" page and keep it with your notes. If you later serve straight from the Hub, the--revisionflag accepts a branch, tag or commit id, so you can pin it there.Record a checksum manifest for the files you downloaded. It costs seconds and lets you prove later that the weights on the server are the weights you approved.
Install the Hugging Face CLI (Linux or macOS) · bash curl -LsSf https://hf.co/cli/install.sh | bashDownload the model to local disk · bash sudo mkdir -p /srv/models sudo chown "$USER" /srv/models hf download Qwen/Qwen3-8B --local-dir /srv/models/Qwen3-8BRecord a checksum manifest · bash cd /srv/models/Qwen3-8B && sha256sum *.safetensors config.json > MANIFEST.sha256Checked against: Hugging Face CLI guide, vLLM troubleshooting, Qwen3-8B model card, vLLM engine arguments
Step 4Start vLLM in Docker, bound to localhost, with an API key
You end up with: A running container that serves an OpenAI-compatible API on 127.0.0.1:8000 and demands a key.
This command follows the Docker example in the vLLM documentation and adds four things: a pinned image tag, a localhost-only port binding, a restart policy and an API key. The docs' own example uses
latest; on 6 October 2026 the Docker Hub tags forvllm/vllm-openaifollow the patternv0.31.0-cu129, and PyPI lists 0.31.0 as the newest release. Check the tags page when you run this and pin whichever version you intend to test.The
-p 127.0.0.1:8000:8000binding matters. Docker's documentation says ports published without an address are published on all host addresses, so a bare-p 8000:8000can expose the API beyond the host. Binding to 127.0.0.1 keeps it private until your proxy is ready.--ipc=hostgives the container the shared memory that tensor-parallel inference uses.--restart unless-stoppedbrings the server back after a reboot unless you stopped it on purpose.--served-model-namesets the name clients must send, so you can swap the underlying model later without changing every client.--max-model-lencaps context length; set it to what you actually need, because it directly limits cache use.--api-keymakes the server require a key, but only on the /v1, /v2 and /inference paths. Step 6 deals with the rest.Generate a key and keep it somewhere safe · bash export VLLM_KEY="$(openssl rand -hex 32)" echo "$VLLM_KEY"Run vLLM (check the tag on Docker Hub first) · bash docker run -d --name vllm --restart unless-stopped \ --runtime nvidia --gpus all \ -v /srv/models:/models \ -p 127.0.0.1:8000:8000 \ --ipc=host \ vllm/vllm-openai:v0.31.0-cu129 \ --model /models/Qwen3-8B \ --served-model-name qwen3-8b \ --max-model-len 16384 \ --api-key "$VLLM_KEY"Without Docker: install into a virtual environment and serve (NVIDIA CUDA) · bash uv venv --python 3.12 --seed source .venv/bin/activate uv pip install vllm --torch-backend=auto vllm serve /srv/models/Qwen3-8B --served-model-name qwen3-8b --max-model-len 16384 --api-key "$VLLM_KEY"Watch it start · bash docker logs -f vllmChecked against: vLLM Docker deployment, vLLM quickstart, vllm serve CLI reference, Docker restart policies, Docker port publishing, vllm/vllm-openai tags on Docker Hub, vLLM on PyPI
Step 5Test the API from the server itself
You end up with: A chat completion comes back, and a request without the key is refused.
Query the server from the host with curl before anything else touches it. The model list confirms the served name; the chat request confirms the model actually generates. The vLLM quickstart shows the same two calls for the chat and completions endpoints on port 8000.
Then send a request without the Authorization header. The /v1 paths should refuse it. If they do not, the key was not applied and you must fix that before the next step.
List models (authenticated) · bash curl -s http://127.0.0.1:8000/v1/models -H "Authorization: Bearer $VLLM_KEY"Ask a question · bash curl -s http://127.0.0.1:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $VLLM_KEY" \ -d '{"model": "qwen3-8b", "messages": [{"role": "user", "content": "Reply with the single word: ready"}], "max_tokens": 20}'Confirm a request with no key is refused · bash curl -i http://127.0.0.1:8000/v1/modelsChecked against: vLLM quickstart, vLLM online serving APIs
Step 6Put a TLS reverse proxy in front and expose only /v1
You end up with: Clients reach the model over HTTPS, and nothing except /v1 is reachable from outside.
The vLLM security page is direct about this. The
--api-keyflag only covers paths under /v1, /v2, /inference and /cohere, and many other sensitive endpoints sit on the same server without authentication. It recommends a reverse proxy such as nginx, Envoy or a Kubernetes Gateway that allowlists only the endpoints you want to expose and blocks the rest. It also says nodes should sit on a dedicated, isolated network with all incoming connections blocked except the API port.Caddy does this in a few lines and obtains certificates automatically for a public domain name. The Caddyfile below sends /v1 traffic to vLLM and returns 404 for everything else, including /metrics and /health.
flush_interval -1switches off response buffering so streamed tokens reach the client as they are produced.If the server has no public DNS name, automatic certificates will not work. Either use an internal certificate authority with your proxy, or give vLLM its own certificate with
--ssl-certfileand--ssl-keyfileand keep it on a private network. Either way, open the firewall to the proxy port only.Caddyfile (replace the domain) · text llm.example.com { @api path /v1/* handle @api { reverse_proxy localhost:8000 { flush_interval -1 } } handle { respond "Not found" 404 } }Run Caddy from the directory that holds the Caddyfile · bash caddy runCheck from another machine · bash curl -s https://llm.example.com/v1/models -H "Authorization: Bearer $VLLM_KEY" curl -i https://llm.example.com/metricsThe second request in the last block should return
HTTP/2 404Checked against: vLLM security guide, Caddy reverse proxy quick start, Caddyfile matchers, Caddy reverse_proxy directive, vllm serve CLI reference
Step 7Add health checks and metrics
You end up with: You can tell whether the server is alive and whether it is running out of cache, before users complain.
From the host,
/healthanswers whether the server is up and/metricsreturns Prometheus-format metrics. Both are blocked at the proxy by design, so scrape them from inside the host or the private network.Four metrics tell you most of what you need.
vllm:num_requests_runningandvllm:num_requests_waitingshow load and queueing.vllm:kv_cache_usage_percshows how full the cache is; when it sits near the top, requests get preempted and latency climbs.vllm:time_to_first_token_secondsis a histogram of how long users wait for the first word. Alert on a growing waiting queue and on cache usage staying high, rather than on a single spike.Then load it. The vLLM benchmark command sends a realistic request mix and reports time to first token, time per output token, inter-token latency and throughput. Run it with a dataset that looks like your traffic, not a toy one, and keep the output with your sizing notes.
Liveness and metrics, from the host · bash curl -s -o /dev/null -w "%{http_code}\n" http://127.0.0.1:8000/health curl -s http://127.0.0.1:8000/metrics | grep -E "vllm:(num_requests_running|num_requests_waiting|kv_cache_usage_perc)"Benchmark the running server (dataset path is yours to supply) · bash vllm bench serve \ --backend vllm \ --model qwen3-8b \ --endpoint /v1/completions \ --dataset-name sharegpt \ --dataset-path /path/to/ShareGPT_V3_unfiltered_cleaned_split.json \ --num-prompts 200Checked against: vLLM online serving APIs, vLLM metrics design, vLLM benchmark CLI
Step 8Tune for your load, and add GPUs only when the numbers say so
You end up with: A short list of settings you have changed on purpose, each with a measured reason.
When the cache is too small for the load, vLLM preempts and recomputes requests, which raises latency. Its optimisation guide lists four levers: raise
--gpu-memory-utilizationto give the cache more room; lower the maximum number of concurrent sequences or batched tokens to need less cache; raise--tensor-parallel-sizeto shard the weights across GPUs so each has more memory left for cache; or use pipeline parallelism to spread layers across GPUs. More parallelism adds synchronisation cost, so it is not free.Change one thing at a time and rerun the same benchmark. Lowering
--max-model-lenis the cheapest lever if your users do not need long contexts. Moving to a quantised checkpoint is the next one, but evaluate the quantised model on your own tasks first, because quality and refusal behaviour can shift.Add a second GPU when a single card cannot hold your weights and cache at the context you need, not because the first benchmark looked slow. Two cards on one machine use
--tensor-parallel-size 2. Multi-node serving is a different project, and the security page warns that traffic between nodes is unencrypted by default and must stay on an isolated network.Example: the same server across two GPUs · bash docker run -d --name vllm --restart unless-stopped \ --runtime nvidia --gpus all \ -v /srv/models:/models \ -p 127.0.0.1:8000:8000 \ --ipc=host \ vllm/vllm-openai:v0.31.0-cu129 \ --model /models/Qwen3-8B \ --served-model-name qwen3-8b \ --tensor-parallel-size 2 \ --api-key "$VLLM_KEY"Checked against: vLLM optimization and tuning, vLLM engine arguments, vLLM security guide
Step 9Plan upgrades and rollback before you need them
You end up with: A repeatable routine: new version beside the old, tested on your prompts, switched by a config change.
vLLM ships often; PyPI shows 0.31.0 released on 5 October 2026. Every new version, new image tag, new quantisation and new model revision can change behaviour. The troubleshooting page even has a flag for one such case:
--generation-config vllmreverts generation defaults to vLLM's neutral ones when quality seems to have changed after switching models.Run the candidate on a second port, say 8001, with the same flags and a different container name. Replay a fixed set of your own prompts against old and new, compare answers and latency, and only then change the proxy to point at the new port. Keep the old container stopped but not deleted for a few days. Rolling back is then one line in the Caddyfile and
caddy reload.Write each change down: image tag, model revision, flags, date, who approved it. That record is also what you will want if anyone asks what was running on a given day. For a fuller regression routine, read how to validate your AI.
Start the candidate beside the current server · bash docker run -d --name vllm-next --runtime nvidia --gpus all \ -v /srv/models:/models \ -p 127.0.0.1:8001:8000 \ --ipc=host \ vllm/vllm-openai:NEW_TAG \ --model /models/Qwen3-8B \ --served-model-name qwen3-8b \ --max-model-len 16384 \ --api-key "$VLLM_KEY"Checked against: vLLM troubleshooting, vLLM on PyPI, vllm/vllm-openai tags on Docker Hub
Is it cheaper than an API? A calculation, not a verdict
Self-hosting turns a per-token bill into a fixed monthly cost, whether the GPU is busy or idle. The comparison has two numbers. First, your monthly GPU cost: the rental or amortised purchase price, plus power, hosting and the share of an engineer's time you will spend. Second, the monthly spend you would have on a hosted API for a comparable model, which you get from your token counts and the provider's current price page.
Self-hosting wins when your usage is high and steady, because the GPU is busy most of the day. It loses when usage is spiky, because you pay for idle hours. We are not giving a break-even figure here because it depends on prices that change; use our token cost calculator and API pricing pages to fill in the API side, and your own quotes for the GPU side.
Count the people. Upgrades, security patches, monitoring and on-call are permanent work, not a one-off setup.
When to stop and use something else
- Your usage is small or irregular: use a hosted API behind a gateway, and revisit when volume grows.
- You need the strongest available model for hard reasoning: open-weight models may trail the best hosted ones on some tasks, so test on your own work first.
- Nobody on the team can own on-call for a GPU server: use a managed dedicated deployment instead of building the operations function yourself.
- Your data must stay in a facility that has no outbound internet: this guide still applies, but plan for an air-gapped environment with offline image and weight transfer.
Troubleshooting
| What you see | Likely cause | Fix |
|---|---|---|
| The container exits during startup with an out-of-memory error | Weights plus the KV cache for your --max-model-len do not fit in the share of GPU memory vLLM may use. | Lower --max-model-len, use a quantised or smaller model, or add a GPU with --tensor-parallel-size. The vLLM troubleshooting page points to its memory-conservation options for this. |
| Docker starts the container but vLLM reports no GPU, or Docker cannot find a GPU runtime | The NVIDIA Container Toolkit is not registered with Docker, or the driver is missing. | Run the nvidia-ctk runtime configure and systemctl restart docker commands from step 2 and re-test with the nvidia-smi sample container before touching vLLM again. |
| Startup hangs while downloading the model | The in-process download stalls, often on a slow or restricted network. | Download the weights first with hf download, then pass the local path as --model. This is the fix the vLLM troubleshooting page gives. |
| Startup is very slow while loading weights from disk | The weights live on a shared or network filesystem. | Copy them to a local disk. The troubleshooting page names network filesystems as a cause of hangs when loading from disk. |
An error appears near self.graph.replay() | A problem with CUDA graph capture on your setup. | Add --enforce-eager to disable CUDA graphs. Expect lower speed, so treat it as a diagnostic first. |
| Requests succeed from curl without the key, or /metrics is reachable from the internet | The key only protects /v1, /v2 and /inference, and the port is published on all addresses. | Bind with -p 127.0.0.1:8000:8000, front the server with the Caddy allowlist from step 6 and block direct access to port 8000 at the firewall. |
| Answers read differently after you switch models or versions | Generation defaults taken from the model's own config differ from the ones you tested. | Start with --generation-config vllm to get vLLM's neutral defaults, then set sampling parameters explicitly in your requests. |
Latency climbs and vllm:num_requests_waiting keeps growing | Demand exceeds the cache or compute you provisioned, so requests queue and may be preempted. | Apply the levers in step 8 one at a time, re-benchmark after each, and add capacity only if the settings cannot fix it. |
Verify it worked
Next steps
- How to set up an LLM gateway: per-team keys, budgets and logs in front of the endpoint
- How to evaluate an open-source LLM: choose and test the model before you size the server
- How to monitor AI agents in production: go beyond GPU metrics to what the model is doing
- Self-hosted LLM inference on Swfte: the engineering detail behind this setup
Related guides
- How to Run LLMs Locally: Ollama, LM Studio, llama.cpp: Install Ollama, LM Studio or llama.cpp, download a model that fits your memory, chat with it and call it from code through a local OpenAI-compatible endpoint.
- How to Set Up an LLM Gateway with LiteLLM (2026): Run the open-source LiteLLM proxy in Docker with a config file, add a model, issue virtual keys with budgets, set fallbacks, wire health checks, and harden it for production.
- How to Evaluate an Open-Source LLM: Hands-On Steps: Check the licence first, write a small task set of your own, run it against the full-precision and quantised model, add a public benchmark as a sanity check, measure speed and memory, and write down the decision.
- How to Build an Air-Gapped AI Environment (2026): Stage model weights, container images and Python packages on a connected machine, verify and carry them across, run the model with every online lookup switched off, and prove nothing leaves.
- How to Self-Host a ChatGPT Alternative (Open WebUI): A hands-on setup of Open WebUI in front of a model you host, with admin accounts, roles, HTTPS, backups and a clear view of where prompts go.
Frequently asked questions
What is vLLM?
vLLM is an open-source inference and serving engine for large language models, licensed Apache-2.0. It runs a server that speaks the OpenAI API format and is built to serve many concurrent requests efficiently on a shared GPU.
Is vLLM better than Ollama?
They suit different jobs. vLLM targets shared servers with many users. Ollama targets simple local use on a laptop or workstation. Many teams prototype on Ollama and move shared workloads to vLLM once they can measure demand.
How much GPU memory do I need to self-host an LLM?
Weights need about two bytes per parameter in 16-bit, one in 8-bit and half a byte in 4-bit, plus a KV cache that grows with context length and concurrent users. An 8.2-billion-parameter model needs about 16.4 GB for 16-bit weights alone, so a 24 GB card leaves limited cache room.
Can I self-host an LLM without a GPU?
Small models run on CPUs and on Apple Silicon through tools such as Ollama or llama.cpp, slowly. vLLM's documented GPU path needs an NVIDIA card with compute capability 7.5 or higher. For shared use with acceptable speed, plan for a GPU.
Is it cheaper to self-host an LLM than to use an API?
Only at high, steady volume. A GPU costs the same whether it is busy or idle, while an API charges per token. Add power, hosting and engineering time to the GPU side before comparing. Below steady heavy use, an API is usually cheaper.
Is Text Generation Inference deprecated?
Its README says TGI is in maintenance mode and will take only minor fixes, documentation improvements and lightweight maintenance, and recommends vLLM, SGLang, llama.cpp or MLX for new work. Existing deployments keep running, but start new ones elsewhere.
Does the vLLM API key protect the whole server?
No. The documentation says --api-key authenticates only paths under /v1, /v2 and /inference, and other endpoints on the same server are not covered. Put a reverse proxy in front that allows only the paths you want to expose.
How Swfte can help
You can complete every step above without Swfte. If you would rather not run the GPU server yourself, Swfte offers dedicated infrastructure and a model gateway.
- Dedicated cloud: isolated infrastructure that Swfte operates for you
- Swfte Connect (model gateway): one OpenAI-compatible endpoint in front of your self-hosted and hosted models
- Deploy models: the pick, evaluate, harden, deploy and monitor path
Swfte Connect is designed to run in your own cloud or data centre. Availability, versions and install details for a customer-hosted deployment: <self-hosted Connect availability - founder to fill>.
Missing a step or found a command that no longer works? Tell us, or request a how-to.
Sources and last verified
Commands, versions and facts in this guide were checked against the sources below on . Tools change quickly: if something differs from what you see, trust the official documentation and let us know.
- vLLM quickstart: uv install commands, vllm serve, default port 8000 curl examples, Linux and Python 3.10 to 3.13
- vLLM Docker deployment: docker run with --runtime nvidia --gpus all, HF cache mount, --ipc=host, --shm-size alternative
- vLLM GPU installation: compute capability 7.5 or higher, CUDA 12.9 binaries
- vllm serve CLI reference: --api-key behaviour, --port default 8000, --ssl-certfile and --ssl-keyfile
- vLLM engine arguments: --gpu-memory-utilization default 0.92, --max-model-len, --tensor-parallel-size, --served-model-name, --revision
- vLLM online serving APIs: /v1/chat/completions, /v1/models, /health and /metrics endpoints
- vLLM security guide: API key path coverage limits, reverse proxy allowlist, isolated network, inter-node traffic
- vLLM metrics design: metric names such as vllm:num_requests_running and vllm:kv_cache_usage_perc
- vLLM optimization and tuning: preemption causes and the four mitigation options
- vLLM troubleshooting: download hang fix, local disk advice, --enforce-eager, --generation-config vllm
- vLLM benchmark CLI: vllm bench serve example and reported metrics
- vLLM on PyPI: version 0.31.0 released 5 October 2026, Apache-2.0, Python 3.10 or newer
- Text Generation Inference README: maintenance-mode notice, recommended alternatives, Apache-2.0
- Ollama FAQ: default bind 127.0.0.1:11434, OLLAMA_NUM_PARALLEL and OLLAMA_MAX_LOADED_MODELS
- NVIDIA Container Toolkit install guide: nvidia-ctk runtime configure, restarting Docker, nvidia-smi sample workload, driver prerequisite
- Docker restart policies: --restart unless-stopped behaviour
- Docker port publishing: -p 127.0.0.1:8080:80 syntax and default all-address binding
- Hugging Face CLI guide: install script, hf download with --local-dir
- Qwen3-8B model card: apache-2.0 licence, 8.2B parameters, context length
- Qwen3-8B config.json: 36 layers, 8 key-value heads, head dimension 128, bfloat16
- Caddy reverse proxy quick start: Caddyfile reverse_proxy, caddy run, automatic HTTPS
- Caddyfile matchers: named path matchers, handle blocks, respond 404 fallback
- Caddy reverse_proxy directive: flush_interval -1 for unbuffered streaming
Topics
- vllm
- self-hosting
- gpu sizing
- openai-compatible api
- inference
Machine-readable copies: this guide as markdown, index of all guides (JSON). Canonical address: https://www.swfte.com/how-to-self-host-an-llm.