Deployment guide
Self-hosted LLM inference: engines, batching, memory and cost
The engineering of running inference yourself, and a method for judging whether it pays.
Self-hosting inference trades a per-token invoice for a fixed fleet you operate. Whether that is a good trade depends on throughput per GPU, how busy you keep the fleet, and what the people who run it cost. This guide covers how modern inference engines get throughput, the memory that limits them, the metrics that tell you they are healthy, and a method for the cost comparison that uses your own numbers rather than anyone’s marketing.
Where the throughput comes from
Continuous batching
Requests join and leave the running batch at every step instead of waiting for a whole batch to finish, which keeps the GPU busy when requests have different lengths.
Paged KV cache
Attention keys and values are stored in fixed-size blocks that are allocated on demand, so memory is not wasted reserving the maximum length for every sequence.
Prefix caching
Requests that share a prefix, such as a long system prompt, reuse the cached blocks instead of recomputing them. It matters most for agents and retrieval with stable preambles.
Parallelism
Tensor parallelism splits layers across GPUs for models too big for one card. Pipeline, expert and data parallelism cover multi-node and mixture-of-experts cases.
Quantization
Lower-precision weights (FP8, AWQ, GPTQ, 4-bit) cut memory and can raise throughput. Re-run your evaluation on the quantized artifact.
Chunked prefill
Long prompts are processed in pieces so they do not stall decoding for everyone else, smoothing time to first token.
Memory is the constraint to plan around
Inference has two phases with different bottlenecks. Prefill, which processes the prompt, is compute-bound. Decode, which generates one token at a time, is memory-bandwidth-bound: each token needs the weights and the KV cache read from GPU memory. That is why memory bandwidth, not just peak compute, predicts generation speed, and why our GPU reference lists bandwidth next to memory size.
KV cache grows with context length times the number of concurrent sequences. For short contexts it is small next to the weights; for long contexts and large batches it can dominate. Set the maximum model length to what your application needs rather than the model’s maximum, because engines reserve memory against it.
For mixture-of-experts models, remember that sparsity saves compute, not memory. Size the cluster for total parameters, then expect speed closer to the active-parameter count.
What to measure and why
| Signal | Why it matters | Typical action |
|---|---|---|
| Time to first token | What users feel as responsiveness; rises with queueing and long prompts. | Add replicas, enable chunked prefill, cap prompt length. |
| Inter-token latency / tokens per second | Streaming smoothness; drops as batch size and context grow. | Reduce max batch, add parallelism or a faster card. |
| Queue depth | Leading indicator of saturation before latency degrades. | Scale out on this, not on CPU. |
| KV-cache utilization | When it nears full, requests are preempted or rejected. | Lower max model length, add memory, quantize the cache. |
| Cost per completed task | Tokens per task vary by model; a verbose model costs more per task at the same price per token. | Compare models on task cost, not per-million-token price. |
| Evaluation drift | Quality and safety can change after a quantization or engine upgrade. | Re-run the suite; roll back on regression. |
vLLM exposes Prometheus-compatible metrics at /metrics and a health check at /health, which is enough to wire autoscaling and alerting.
A cost method that uses your numbers
Compute the fleet cost per month: GPUs times hourly price (or amortized hardware plus power and hosting) times hours kept up. Estimate sustained throughput in tokens per second per replica on your model, prompt shape and concurrency by measuring it; do not use someone else’s benchmark. Multiply by the utilization you realistically achieve, because a fleet sized for peak sits partly idle.
Divide cost by tokens actually served to get an effective price per million tokens, then add the people: the engineer who keeps it alive and the on-call rotation. Compare that with the hosted price for the same quality. Where traffic is low or spiky, hosted is usually cheaper. Self-hosting pays when volume is high and steady, when prompts cannot leave your perimeter, when you need to fine-tune, or when vendor independence is worth a premium. Those last reasons are not about price, so decide what you are buying.
We do not publish a single break-even figure here, because it depends on your model, hardware, utilization and staffing, and an invented one would mislead.
Autoscaling and routing in practice
Scale replicas on queue depth and time to first token. Keep a minimum warm count for latency-sensitive routes because weight loading is slow, and scale to zero only for batch traffic that tolerates a cold start.
Put the Connect gateway in front so routing is policy, not code: send high-volume, low-risk work to the cheaper model, keep the harder slice on a stronger one, and fail over to another approved model when a replica set is unhealthy.
Frequently asked questions
What is continuous batching?
It adds new requests to the running batch and removes finished ones at every generation step, instead of waiting for a whole batch to finish. It keeps the GPU busy when requests differ in length.
Why does memory bandwidth matter for LLM inference?
During decode, every generated token needs the model weights and KV cache read from GPU memory, so generation speed is often limited by memory bandwidth rather than peak compute.
When is self-hosting cheaper than a hosted API?
When volume is high and steady and you can keep GPUs busy. At low or spiky volume a hosted API is usually cheaper once engineering and on-call are counted. Measure your own throughput and utilization to decide.
What should I autoscale on?
Queue depth and time to first token, not CPU. Keep warm replicas for latency-sensitive routes.
Does a quantized KV cache change quality?
It can. Treat any change to precision as a new candidate and re-run your evaluation suite.