Deployment guide

Self-hosted LLM inference: engines, batching, memory and cost

The engineering of running inference yourself, and a method for judging whether it pays.

Self-hosting inference trades a per-token invoice for a fixed fleet you operate. Whether that is a good trade depends on throughput per GPU, how busy you keep the fleet, and what the people who run it cost. This guide covers how modern inference engines get throughput, the memory that limits them, the metrics that tell you they are healthy, and a method for the cost comparison that uses your own numbers rather than anyone’s marketing.

Where the throughput comes from

Continuous batching

Requests join and leave the running batch at every step instead of waiting for a whole batch to finish, which keeps the GPU busy when requests have different lengths.

Paged KV cache

Attention keys and values are stored in fixed-size blocks that are allocated on demand, so memory is not wasted reserving the maximum length for every sequence.

Prefix caching

Requests that share a prefix, such as a long system prompt, reuse the cached blocks instead of recomputing them. It matters most for agents and retrieval with stable preambles.

Parallelism

Tensor parallelism splits layers across GPUs for models too big for one card. Pipeline, expert and data parallelism cover multi-node and mixture-of-experts cases.

Quantization

Lower-precision weights (FP8, AWQ, GPTQ, 4-bit) cut memory and can raise throughput. Re-run your evaluation on the quantized artifact.

Chunked prefill

Long prompts are processed in pieces so they do not stall decoding for everyone else, smoothing time to first token.

Memory is the constraint to plan around

Inference has two phases with different bottlenecks. Prefill, which processes the prompt, is compute-bound. Decode, which generates one token at a time, is memory-bandwidth-bound: each token needs the weights and the KV cache read from GPU memory. That is why memory bandwidth, not just peak compute, predicts generation speed, and why our GPU reference lists bandwidth next to memory size.

KV cache grows with context length times the number of concurrent sequences. For short contexts it is small next to the weights; for long contexts and large batches it can dominate. Set the maximum model length to what your application needs rather than the model’s maximum, because engines reserve memory against it.

For mixture-of-experts models, remember that sparsity saves compute, not memory. Size the cluster for total parameters, then expect speed closer to the active-parameter count.

What to measure and why

SignalWhy it mattersTypical action
Time to first tokenWhat users feel as responsiveness; rises with queueing and long prompts.Add replicas, enable chunked prefill, cap prompt length.
Inter-token latency / tokens per secondStreaming smoothness; drops as batch size and context grow.Reduce max batch, add parallelism or a faster card.
Queue depthLeading indicator of saturation before latency degrades.Scale out on this, not on CPU.
KV-cache utilizationWhen it nears full, requests are preempted or rejected.Lower max model length, add memory, quantize the cache.
Cost per completed taskTokens per task vary by model; a verbose model costs more per task at the same price per token.Compare models on task cost, not per-million-token price.
Evaluation driftQuality and safety can change after a quantization or engine upgrade.Re-run the suite; roll back on regression.

vLLM exposes Prometheus-compatible metrics at /metrics and a health check at /health, which is enough to wire autoscaling and alerting.

A cost method that uses your numbers

Compute the fleet cost per month: GPUs times hourly price (or amortized hardware plus power and hosting) times hours kept up. Estimate sustained throughput in tokens per second per replica on your model, prompt shape and concurrency by measuring it; do not use someone else’s benchmark. Multiply by the utilization you realistically achieve, because a fleet sized for peak sits partly idle.

Divide cost by tokens actually served to get an effective price per million tokens, then add the people: the engineer who keeps it alive and the on-call rotation. Compare that with the hosted price for the same quality. Where traffic is low or spiky, hosted is usually cheaper. Self-hosting pays when volume is high and steady, when prompts cannot leave your perimeter, when you need to fine-tune, or when vendor independence is worth a premium. Those last reasons are not about price, so decide what you are buying.

We do not publish a single break-even figure here, because it depends on your model, hardware, utilization and staffing, and an invented one would mislead.

Autoscaling and routing in practice

Scale replicas on queue depth and time to first token. Keep a minimum warm count for latency-sensitive routes because weight loading is slow, and scale to zero only for batch traffic that tolerates a cold start.

Put the Connect gateway in front so routing is policy, not code: send high-volume, low-risk work to the cheaper model, keep the harder slice on a stronger one, and fail over to another approved model when a replica set is unhealthy.

Frequently asked questions

What is continuous batching?

It adds new requests to the running batch and removes finished ones at every generation step, instead of waiting for a whole batch to finish. It keeps the GPU busy when requests differ in length.

Why does memory bandwidth matter for LLM inference?

During decode, every generated token needs the model weights and KV cache read from GPU memory, so generation speed is often limited by memory bandwidth rather than peak compute.

When is self-hosting cheaper than a hosted API?

When volume is high and steady and you can keep GPUs busy. At low or spiky volume a hosted API is usually cheaper once engineering and on-call are counted. Measure your own throughput and utilization to decide.

What should I autoscale on?

Queue depth and time to first token, not CPU. Keep warm replicas for latency-sensitive routes.

Does a quantized KV cache change quality?

It can. Treat any change to precision as a new candidate and re-run your evaluation suite.

Run inference on dedicated GPUs with Swfte

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.