LLM runtimes and inference servers
compared.
These tools split cleanly into two groups that are often compared but rarely interchangeable: runners built to get a model working on one machine, and servers built to keep a GPU saturated under concurrent traffic. Choosing across the boundary is the most common mistake.
published directory
Two groups.
Rarely interchangeable.
Local runners are optimised for getting a model running on a single machine with minimal setup. No continuous batching, so throughput under parallel load is limited by design — that is a trade, not a defect.
Inference servers are built for serving. Continuous batching and paged attention keep the GPU busy across many concurrent requests, and tensor parallelism lets a model span several cards. You own the process and its operations.
6 entries / Page 1 of 1
Ollama
Inspect runtime ↗local runner / openai compatible api / runs on cpuLM Studio
Inspect runtime ↗inference server / continuous batching / paged attention / tensor parallel / openai compatible apivLLM
Inspect runtime ↗local runner / inference server / continuous batching / openai compatible api / runs on cpullama.cpp
Inspect runtime ↗inference server / continuous batching / paged attention / tensor parallel / openai compatible apiSGLang
Inspect runtime ↗inference server / continuous batching / paged attention / tensor parallel / openai compatible apiText Generation Inference (TGI)
Inspect runtime ↗Read them
head-to-head.
The pairs people actually weigh against each other, each one read on batching, memory management, quantisation and the deployment model it assumes.
Ollama vs vLLM vLLM vs SGLang Ollama vs LM Studio vLLM vs Text Generation Inference (TGI) Ollama vs llama.cppA runtime is
an operating choice.
Decide what the process has to serve, who operates it, and where its weights and traffic are allowed to live — before the model is chosen.
Run models on dedicated capacity Define access boundaries