Self-Hosted LLM Stack: vLLM, Ollama, LiteLLM and More
The layers of a self-hosted LLM stack and the open-source tools for each, from vLLM and Ollama to LiteLLM.
A self-hosted LLM stack has six layers: hardware, an inference engine, a gateway, retrieval, a user-facing application and a governance layer that records and limits what happens. vLLM and Ollama sit at the inference layer, a gateway such as LiteLLM sits above them, and the rest is yours to assemble or buy. The most common mistake is to treat the inference engine as the whole stack and discover the other five layers in production.
This guide walks through each layer, names the open-source options that are widely used and states what each is for. The aim is a map you can use to decide what to build, what to adopt and what to hand to a platform.
What are the layers of a self-hosted AI platform?
| Layer | Job | Typical open-source options |
|---|---|---|
| Hardware | Run the model | GPUs, or CPUs and Apple Silicon for small models |
| Inference engine | Load weights, serve tokens | vLLM, Ollama, others |
| Gateway | One endpoint, keys, routing, spend | LiteLLM and similar |
| Retrieval | Embeddings, vector store, ingestion | Embedding models plus a vector database |
| Application | Chat, workspace, agents | Open WebUI, LibreChat, AnythingLLM |
| Governance | Identity, policy, audit, evaluation | Mostly custom or platform-provided |
A self-hosted AI platform is these six layers working together. You can host all of them, host some and use managed services for others, or buy the integrated whole.
Which inference engine: vLLM or Ollama?
They solve different problems, and many teams use both.
vLLM is a serving engine for production throughput. It is Apache-2.0 licensed, runs on NVIDIA, AMD and other hardware, exposes an OpenAI-compatible API server and uses PagedAttention and continuous batching to serve many concurrent requests, according to its repository. It supports a long list of quantization formats and parallelism strategies. Use it when many users share a server and you care about tokens per second per dollar. Our continuous batching deep dive explains why batching matters so much, and the serving frameworks comparison covers alternatives.
Ollama is a tool for running open models locally with minimal setup. It is MIT licensed, ships a command-line tool and a REST API, uses llama.cpp as a backend, runs on macOS, Windows and Linux, and serves a local API on port 11434 by default, per its repository. Use it for workstations, prototypes and small teams. It is not designed to maximize multi-user throughput on a GPU server.
| Question | vLLM | Ollama |
|---|---|---|
| Primary use | Shared, high-throughput serving | Local and small-team use |
| Setup effort | Moderate | Low |
| Concurrency focus | Yes | Limited |
| Runs on a laptop | Not the target | Yes |
| API | OpenAI-compatible | Own REST API, plus compatibility layers |
A sensible path: prototype on Ollama, then move shared workloads to vLLM when you can measure demand.
Why put a gateway in front?
Without a gateway, every application embeds a model address and, often, a provider key. Swapping a model means touching every application, and nobody can say what the usage cost.
A gateway gives you one endpoint, central keys, routing, fallbacks and spend tracking. LiteLLM is a popular open-source option: its proxy server presents an OpenAI-compatible interface across 100+ model providers including Ollama and vLLM, supports virtual keys and spend tracking, and offers caching, retries and fallbacks, per its repository. It also has a commercial enterprise tier, so check which features you need fall in which tier. It starts a local server on port 4000 in its quick start.
The effect is that applications talk to one address:
from openai import OpenAI
# Applications only know the gateway, never the model server.
client = OpenAI(base_url="http://llm-gateway.internal:4000", api_key="team-virtual-key")
reply = client.chat.completions.create(
model="internal-default", # an alias the gateway maps to a model
messages=[{"role": "user", "content": "Summarize this incident report."}],
)
When you change the model behind internal-default, nothing in the application changes. That is how you avoid lock-in in practice, and the vendor lock-in guide and model exit-cost audit show why it matters. The AI gateways post discusses the build-versus-buy trade-off.
The gateway is also your enforcement point for data classes. Route restricted traffic only to local models and block it from hosted endpoints. Without that rule in the gateway, privacy relies on user behavior.
Which models should a self-hosted stack run?
The open-weight field changes monthly, so choose by license, size and task rather than by leaderboard alone. Ollama's own README lists families such as Kimi, GLM, MiniMax, DeepSeek, gpt-oss, Qwen and Gemma as runnable, which gives a sense of the breadth. OpenAI's gpt-oss models are released under Apache 2.0, with the larger one designed to fit on a single 80 GB GPU.
Practical rules:
- Read the license. Some open-weight licenses restrict commercial use or large deployments.
- Do the memory arithmetic. Weights need about two bytes per parameter at 16-bit, one at 8-bit and half a byte at 4-bit, plus room for the KV cache.
- Evaluate on your own tasks. Public benchmarks do not predict your documents.
- Keep two models. A small, fast model for routine tasks and a larger one for hard ones, selected through the gateway.
Our posts on open weights versus proprietary models, Gemma 4 sizing and running Qwen locally go deeper on specific families.
What about retrieval?
Retrieval is how the stack answers from your documents. It needs an embedding model, a vector store and an ingestion pipeline that handles parsing, chunking and refresh. The part that decides whether it is safe is permissions: each retrieved chunk must carry the access rules of its source, and the query must be filtered by the asking user's rights. See the RAG architecture guide and what a knowledge base is.
What does governance look like in a self-hosted stack?
This is the layer open-source stacks leave to you. Plan for:
- Single sign-on and role-based access to models and knowledge.
- Logs of requests, retrieved sources and tool calls, stored where you control them.
- Evaluation: regression tests for model and prompt changes, since a silent model update can change behavior.
- Agent policy: scopes, approval thresholds and rate limits.
- Cost attribution by team.
- Observability. See LLM observability and prompt analytics.
What does it cost to run?
Count four things: hardware or rented GPUs, power and hosting, engineering time and the opportunity cost of delay. The TCO analysis and on-premises economics post give you frameworks for the sums. Utilization drives everything: an idle GPU is the expensive failure mode.
Where does Swfte fit?
If you want the benefits of the stack without assembling every layer, Swfte provides several of them as one platform. BuildX is the gateway across 50+ models, Studio builds agents and workflows, Nexus watches and enforces policy on agent actions, and Cortex is the governed AI desktop that runs local by default. Dedicated cloud provides isolated infrastructure including air-gapped options when you do not want to run GPUs yourself. The infrastructure and governance pages explain how the layers connect.
Swfte is designed to provide technical controls, governance mechanisms and evidence for deploying AI within your own regulatory, security and policy requirements. The exact posture depends on your use case, jurisdiction, deployment and configuration. To plan a stack, contact the team. For the interface layer, read the self-hosted ChatGPT alternative comparison, and for agents, on-prem AI agents.
Frequently asked questions
What is a self-hosted LLM stack?
It is the set of components you run yourself to serve language models and build applications on them: hardware, an inference engine, a gateway, retrieval, an application layer and governance.
Is vLLM better than Ollama?
They suit different jobs. vLLM targets shared, high-throughput serving. Ollama targets simple local use. Many teams prototype on Ollama and move shared workloads to vLLM.
Do I need a gateway like LiteLLM?
Not for a single application. Once several applications or teams use models, a gateway pays for itself through central keys, spend tracking, routing and the ability to swap models without changing applications.
Can a self-hosted stack match hosted frontier models?
For many tasks, open-weight models are good enough. For the hardest tasks, hosted frontier models may still lead, which is why a gateway that can route to both is useful, with data classes deciding what may go where.
How do I keep a self-hosted stack compliant?
No stack is compliant by itself. Provide logging, access control and evidence, classify your AI uses and follow the rules that apply to you. See enterprise AI governance.
Related: Swfte Connect is the model gateway, designed to run in your own cloud or data centre; see self-deploying Connect.