Deployment guide
How to deploy an open-source LLM in production
The end-to-end path from a downloaded checkpoint to a governed production endpoint.
Running an open-source model in production is a systems problem wearing a machine-learning costume. The model is a few hundred gigabytes of numbers; everything that makes it safe and useful is the engineering around it. This guide walks the path we follow: size the hardware, choose the engine, decide on quantization, put an API in front, and refuse to promote anything that has not passed an evaluation gate.
The path, in order
1. Choose the candidate
Pick by task fit, context length, language coverage and licence. Pin the exact repository revision, not a branch name, and keep the licence file with the artifact.
2. Size the memory
Weights in bytes are parameters times bytes per parameter: two for bf16, one for FP8, about half for 4-bit. Add KV cache for your context length and batch size, plus headroom for activations. Mixture-of-experts models need every expert resident.
3. Choose the engine
vLLM for shared GPU serving, SGLang if prefixes repeat, Ollama or llama.cpp for development and edge. Avoid starting new work on Text Generation Inference, which is in maintenance mode.
4. Decide the quantization
Use the highest precision that fits your budget, then treat each quantized artifact as a new candidate that must pass evaluation again.
5. Evaluate, then harden
Run capability, safety and multilingual suites. Verify the weights hash, load safetensors only and switch off remote code. Only then expose it.
6. Deploy behind a gateway
Put the endpoint behind Connect so routing, failover, approved-model policy, logging and cost tracking apply from the first request.
7. Monitor and keep a rollback
Track quality, safety signals, time to first token and cost per completed task. Keep the previous approved revision routable.
Sizing: do the arithmetic before you rent the cluster
Start with weights. A 70-billion-parameter dense model needs about 140 GB in bf16, about 70 GB in FP8 and about 35 GB at 4-bit. A single 80 GB H100 holds the 4-bit version comfortably and the FP8 version barely; the bf16 version needs at least two cards with tensor parallelism. Those are arithmetic results, not measurements, and your real footprint is higher because some tensors (embeddings, normalization, routers) are not quantized the way expert weights are.
Then add KV cache. It grows with context length times concurrent sequences, and for long contexts it can rival the weights. Reduce it by capping the maximum model length at what your application really uses rather than the model’s advertised maximum, and by using an engine with paged KV memory.
Plan with headroom. A configuration that fits on paper with a few gigabytes spare is a demonstration, not a system. Leave room for larger batches and for the next model revision, which is usually bigger.
Choosing the serving engine
| Engine | Best for | Watch out for |
|---|---|---|
| vLLM | Shared GPU serving, OpenAI-compatible API, wide model and quantization support, Prometheus metrics. | Needs a compatible GPU and some operational experience. GGUF support is described by its own documentation as experimental. |
| SGLang | Agent and retrieval workloads with long shared prefixes. | Benchmark it on your traffic before you commit; the advantage depends on prefix reuse. |
| Ollama / llama.cpp | Laptops, workstations, edge, single-team tools. | Built for low concurrency; throughput under many simultaneous users is not its job. |
| TGI | Existing deployments that already work. | In maintenance mode: new architectures and kernels land in other engines first. |
The evaluation gate is what makes it production
A model that answers your demo prompt well has not been tested. Before promotion, run it on a held-out set drawn from your real tasks, on a safety suite covering refusal of harmful requests and over-refusal of benign ones, on prompt-injection cases, and on the languages your users write in. Compare against the version currently in production, not against a leaderboard. The result either clears the gate or it does not.
Do the same for every change that is not a code change: a new quantization, a new serving image, a new system prompt, a new base revision. They all alter behavior. Our methodology page lists the suites and what each one is for.
The API surface and governance around it
Most open-source engines expose an OpenAI-compatible HTTP API, which means existing SDKs and agents work by changing the base URL and key. Do not expose that port directly. Put it on a private network, require authentication, and let the Connect gateway be the only client. The gateway is where you enforce which models a workload may call, log requests, track cost per route and fail over.
Add runtime guardrails at the edges: validate inputs, filter outputs and limit what tools a model-driven agent can call. Guardrails reduce harm at request time; they do not replace the evaluation gate, and the gate does not replace them.
Frequently asked questions
How much GPU memory do I need to deploy an open-source LLM?
Weights take about two bytes per parameter in bf16, one in FP8 and about half a byte at 4-bit, plus KV cache that grows with context length and concurrency, plus headroom. Mixture-of-experts models need all experts in memory even though few are active per token.
Should I use vLLM, SGLang, TGI or Ollama?
vLLM is the sensible default for shared GPU serving. Consider SGLang for prefix-heavy agent workloads. Ollama and llama.cpp suit development and edge. TGI is in maintenance mode, so avoid starting new work on it.
Do I need to re-test after quantizing?
Yes. Quantization changes the artifact you serve, and refusal behavior and non-English quality can shift. Treat the quantized model as a new candidate.
Can I bring my own fine-tuned model?
Yes. A fine-tune of your own goes through the same loop. Fine-tuning can weaken safety behavior even when the data is benign, so the safety suite matters more, not less.
Is the endpoint OpenAI-compatible?
The common engines expose OpenAI-compatible APIs, so existing SDKs usually work by changing the base URL. Place the endpoint behind the Connect gateway rather than exposing it directly.