AI tools / Local and served

LLM runtimes and inference servers
compared.

These tools split cleanly into two groups that are often compared but rarely interchangeable: runners built to get a model working on one machine, and servers built to keep a GPU saturated under concurrent traffic. Choosing across the boundary is the most common mistake.

6runtimes in the
published directory

Two groups.
Rarely interchangeable.

Local runners are optimised for getting a model running on a single machine with minimal setup. No continuous batching, so throughput under parallel load is limited by design — that is a trade, not a defect.

Inference servers are built for serving. Continuous batching and paged attention keep the GPU busy across many concurrent requests, and tensor parallelism lets a model span several cards. You own the process and its operations.

6 entries / Page 1 of 1

Read them
head-to-head.

The pairs people actually weigh against each other, each one read on batching, memory management, quantisation and the deployment model it assumes.

Ollama vs vLLM vLLM vs SGLang Ollama vs LM Studio vLLM vs Text Generation Inference (TGI) Ollama vs llama.cpp

A runtime is
an operating choice.

Decide what the process has to serve, who operates it, and where its weights and traffic are allowed to live — before the model is chosen.

Run models on dedicated capacity Define access boundaries

Build this in Studio

Describe what you need in plain language. Studio builds the agents and workflows, and you keep every version.