Short answer
An LLM gateway is one OpenAI-compatible endpoint in front of your models, so applications hold a gateway key instead of provider keys. The quickest open-source route is the LiteLLM proxy: write a config.yaml with a model_list and a master key, run the Docker image, test with curl, add Postgres to issue virtual keys with budgets, then add fallbacks, health checks and hardening. Pin the version you run.
The steps at a glance
- Choose the first model to put behind the gateway
- Decide how you will start the proxy
- Write config.yaml with one model and a master key
- Run the gateway and send a test request
- Add Postgres and issue virtual keys with budgets
- Set retries, timeouts and fallbacks
- Add health checks and production settings
- Decide what the gateway logs and redact what it should not keep
- Harden the deployment
Before you start
Who this is for
- Platform and backend engineers who want provider keys, budgets and logs in one place instead of in every application.
- Teams running more than one model, or both hosted and self-hosted models, behind one API.
- Anyone replacing scattered provider SDK calls with a single base URL they control.
Probably not for you if
- Teams with one application and one provider and no budget or audit need. A gateway is another service to run, and you may not need it yet.
- Teams that want a hosted service with no infrastructure. Compare the hosted options in the table below first.
Prerequisites
- Docker installed, and permission to run containers on a server or your laptop.
- At least one model to put behind the gateway: a provider API key, or an OpenAI-compatible server such as llama.cpp's
llama-server. - A PostgreSQL database for step 5 onwards (virtual keys and budgets need a database). The LiteLLM quickstart starts one for you with Docker Compose.
- A way to keep secrets out of files you commit: environment variables, a secrets manager, or a
.envfile that stays out of version control.
- Time
- About 2 hours to a working, budgeted gateway; a day to production-harden it
- Cost
- LiteLLM Open Source is free to self-host. You pay your model providers and for the server and database you run it on.
- Hardware
- LiteLLM's production guide recommends 1 vCPU and 4 GiB of memory per pod as a floor. A laptop is enough to follow the steps.
- Skill
- Comfortable with Docker, YAML and curl
Estimates are ours, not measurements, and move with your hardware, data and network.
Run it yourself or call a hosted gateway
The first decision is who carries the pager. LiteLLM is software you operate, so your prompts and keys stay on your infrastructure and so do the uptime and upgrades. A hosted gateway is a service you call, with less to run and a third party in the data path. The facts below come from the vendors' own pages as recorded on our alternatives pages on 6 October 2026. Check them before you buy.
| Option | How you run it | Pricing as published | Worth knowing |
|---|---|---|---|
| LiteLLM | Open-source Python SDK and self-hosted proxy | Open Source: $0, self-hosted. Enterprise: sales, annual | MIT licence except its enterprise directory; virtual keys with budgets and rate limits in the open-source plan |
| OpenRouter | Hosted service | Free plan; pay-as-you-go plans with a fee on credit purchases; inference at provider list price | EU and US in-region routing is listed on the Business plan |
| Portkey | Open-source gateway, hosted platform or private cloud | Open source; Developer free; Production $49 a month; Enterprise custom | Guardrails and semantic caching depend on the plan |
| Helicone | Observability proxy reached by a base-URL change | Open-source core; free tier, then usage-based per logged request | Sees only traffic that passes through the proxy |
| Swfte Connect | Managed gateway on the Swfte platform | Free (pay as you use); Pro 5% platform fee; Scale 3 to 4%; Enterprise negotiable | Dedicated or private deployment is scoped as an engagement, not self-serve |
Step 1Choose the first model to put behind the gateway
You end up with: One model you can already call, with its address and key written down.
Start with a model you can already reach so you can tell a gateway fault from a model fault. Two good choices: a hosted provider where you have an API key, or a local OpenAI-compatible server. llama.cpp's
llama-serveris one:llama-server -m model.gguf --port 8080servesPOST /v1/chat/completionson that port.Write down three things: the provider or server address, the model name that backend expects, and the key (or "none" for a local server with no
--api-key). You will put these in the config file in step 3. Do not put real keys in the file itself; LiteLLM reads them from environment variables, as you will see.If you want to test with no paid account at all, run the local server and use that. Our guide to running LLMs locally covers getting a model onto your machine.
Optional: serve a local model on port 8080 · bash llama-server -m model.gguf --port 8080Checked against: llama.cpp server README
Step 2Decide how you will start the proxy
You end up with: A choice between the one-command quickstart and a config file you control.
LiteLLM documents two starts. The quickstart script deploys the gateway and a Postgres database with Docker Compose, writes generated security keys to a
.envfile, and serves the gateway and an admin UI onhttp://localhost:4000. It is the fastest way to see the product. Read the script before you run it, as you should with anything you pipe into a shell.The second start is a YAML config file and one
docker run, with no database. The documentation describes it as the minimal setup for API-only use and notes that budget enforcement needs a database. This guide uses the config file, because a file you keep in version control is repeatable and reviewable, then adds Postgres in step 5.If you take the quickstart route, you can still follow steps 6 to 9, and you manage models and keys in the admin UI instead of the file.
Option A: the LiteLLM quickstart script (read it first) · bash curl -fsSL https://raw.githubusercontent.com/BerriAI/litellm/main/scripts/quickstart.sh | shChecked against: LiteLLM gateway quickstart, LiteLLM: Security Update, Suspected Supply Chain Incident
Step 3Write config.yaml with one model and a master key
You end up with: A config file naming one model and reading its secrets from environment variables.
The config file has three parts that matter now.
model_listis the list of models clients can ask for. For each entry,model_nameis the name clients use, andlitellm_params.modelis the provider-specific string LiteLLM passes on.general_settings.master_keyis the admin credential, and the documentation shows it read from the environment asos.environ/LITELLM_MASTER_KEY.Model strings use a
provider/modelform, such asanthropic/<model>oropenai/<model>. For a custom OpenAI-compatible server, such as llama.cpp, use theopenai/prefix plus anapi_base. The docs show exactly this pattern for a vLLM endpoint, withapi_key: none. For a hosted provider, put the key in an environment variable and refer to it asos.environ/NAME.Clients never see the provider key. They send the gateway key, and the gateway adds the provider key on the way out. That is most of the value of the pattern.
config.yaml · yaml model_list: - model_name: local-llm litellm_params: model: openai/model api_base: http://host.docker.internal:8080/v1 api_key: none - model_name: hosted-model litellm_params: model: <provider>/<model> api_key: os.environ/PROVIDER_API_KEY litellm_settings: drop_params: True general_settings: master_key: os.environ/LITELLM_MASTER_KEYCreate a master key (it should start with sk-) and keep it out of version control · bash export LITELLM_MASTER_KEY="sk-$(openssl rand -hex 24)" export PROVIDER_API_KEY="<your provider key>"Checked against: LiteLLM config.yaml documentation, LiteLLM providers documentation, LiteLLM production best practices
Step 4Run the gateway and send a test request
You end up with: A completion returned through http://localhost:4000 using your master key.
Run the official image with the config file mounted and the environment variables passed through. The documentation shows the mount form
docker run -v /path/to/config.yaml:/app/config.yaml docker.litellm.ai/berriai/litellm:latest --config /app/config.yaml. Replacelatestwith a version tag you have tested, and publish port 4000.Then send a chat request with curl. The endpoint is
/chat/completionson port 4000, in the OpenAI format. Because you set a master key, send it as a bearer token. Themodelvalue is yourmodel_namefrom the config, not the provider string.Any OpenAI SDK now works by changing two values: the base URL to
http://localhost:4000and the API key to a gateway key. Application code does not otherwise change. That is how you migrate applications onto a gateway one at a time.Run the proxy · bash docker run -p 4000:4000 \ -e LITELLM_MASTER_KEY \ -e PROVIDER_API_KEY \ -v "$(pwd)/config.yaml:/app/config.yaml" \ docker.litellm.ai/berriai/litellm:<tested-version> \ --config /app/config.yamlSend a request through the gateway · bash curl --location 'http://0.0.0.0:4000/chat/completions' \ --header 'Content-Type: application/json' \ --header "Authorization: Bearer $LITELLM_MASTER_KEY" \ --data '{ "model": "local-llm", "messages": [{"role": "user", "content": "Reply with the single word: ready"}] }'Checked against: LiteLLM config.yaml documentation
Step 5Add Postgres and issue virtual keys with budgets
You end up with: Per-application keys, each with its own model list, spend limit and rate limit.
Do not give applications the master key. Create a virtual key for each application or team. LiteLLM generates them with
POST /key/generate, authorised by the master key, and each key can carry amax_budget, abudget_duration, a list of allowedmodels, andrpm_limitandtpm_limitfor requests and tokens per minute.Budgets and virtual keys need a database. The documentation is explicit that budgets require one, and it takes a PostgreSQL connection string in the
DATABASE_URLenvironment variable. Start a Postgres instance (the quickstart Compose file includes one), add-e DATABASE_URLto thedocker runcommand, and restart the proxy.When a key goes over its budget the gateway rejects the request with an error such as
ExceededTokenBudget, which your application must handle like any other API failure. Set budgets slightly above expected use at first and tighten them once you have a month of data. Rate limits per model can also be set inmodel_list, for examplerpm: 6on an entry.Issue a virtual key (replace the values with your own) · bash curl 'http://0.0.0.0:4000/key/generate' \ --header "Authorization: Bearer $LITELLM_MASTER_KEY" \ --header 'Content-Type: application/json' \ --data-raw '{ "max_budget": 10, "budget_duration": "30d", "models": ["local-llm"], "rpm_limit": 60, "tpm_limit": 20000 }'Checked against: LiteLLM virtual keys documentation, LiteLLM config.yaml documentation
Step 6Set retries, timeouts and fallbacks
You end up with: A failing model is retried, then replaced by the next model in your list.
Fallbacks are the feature that makes a gateway earn its place. In
litellm_settingsyou setnum_retries,request_timeout,allowed_failsandcooldown_time, and afallbackslist that maps a model name to the models to try next. The documentation says fallbacks run in order, and that a model which fails more thanallowed_failstimes in a minute is put on cooldown forcooldown_timeseconds.There are specialised fallback lists too:
context_window_fallbacksfor context-length errors,content_policy_fallbacksfor content policy errors anddefault_fallbacksfor anything else. A common pattern is a small, cheap model as the primary, and a larger one as the fallback for context overflows.Test it. Stop the backend for the primary model, send a request, and check the answer arrives from the fallback. A fallback you have not triggered on purpose is a fallback you do not know works. Check the data protection position of every model in the chain: a fallback to a different provider changes where your data goes.
Add to config.yaml · yaml litellm_settings: drop_params: True num_retries: 3 request_timeout: 30 fallbacks: [{"local-llm": ["hosted-model"]}] allowed_fails: 3 cooldown_time: 30Checked against: LiteLLM fallbacks and reliability documentation
Step 7Add health checks and production settings
You end up with: Probes your orchestrator can use, and settings that suit a production deployment.
LiteLLM exposes three health endpoints.
GET /health/livelinesssays the process is up and checks no dependencies.GET /health/readinesssays the worker is ready for traffic, and reports the database state.GET /healthruns a real test request against every configured model, so it costs a few tokens per model: do not point a once-a-second probe at it. Use the first two for container probes and the third for an occasional manual check.The production guide lists the settings to change: export
LITELLM_MODE=PRODUCTIONto stop automatic.envloading, setLITELLM_SALT_KEYto a random value to encrypt stored credentials (and never change it after deployment), write spend updates to the database every 60 seconds withproxy_batch_write_at: 60, keep error logs out of the database withdisable_error_logs: True, and setjson_logs: truefor structured logs. For several instances it recommends Redis 7.0 or later to share rate limits and router state.Put TLS in front of the gateway with your reverse proxy or load balancer, and do not expose the admin UI to the internet. Keep the gateway on a network the applications can reach and the public cannot.
Probe the gateway · bash curl http://localhost:4000/health/liveliness curl http://localhost:4000/health/readinessProduction settings for config.yaml · yaml general_settings: master_key: os.environ/LITELLM_MASTER_KEY proxy_batch_write_at: 60 disable_error_logs: True litellm_settings: json_logs: trueChecked against: LiteLLM health endpoints, LiteLLM production best practices
Step 8Decide what the gateway logs and redact what it should not keep
You end up with: A written logging policy and, where needed, a PII masking guardrail in front of the model.
A gateway sees every prompt and every answer, which is its value and its risk. Decide what you keep. At minimum you want who called, which model, token counts, cost and latency. Whether you also keep prompt and response text depends on your data protection position; keeping them makes debugging easier and makes the log a store of personal data you must protect, retain and delete properly.
LiteLLM supports guardrails, including a Presidio-based PII masking guardrail with a
pre_callmode that masks or blocks entities such as credit card numbers before the request reaches the model. The configuration is aguardrailsblock with aguardrail_nameandlitellm_params. Read the documentation page for the services Presidio needs before you rely on it, and test it on realistic text, because pattern-based masking misses things.Redaction in the gateway does not replace deciding what data should go to a model at all. Combine it with an allow-list of models per key and, for sensitive work, a self-hosted model. For the wider practice, see how to monitor AI agents in production.
Guardrail block from the LiteLLM documentation (needs the Presidio services it describes) · yaml guardrails: - guardrail_name: "presidio-pii" litellm_params: guardrail: presidio mode: "pre_call" presidio_language: "en"Checked against: LiteLLM PII masking guardrail (Presidio)
Step 9Harden the deployment
You end up with: A pinned, minimally exposed gateway with secrets, updates and incident steps decided in advance.
Pin versions. Run an image tag you have tested, not
latest, and promote a new tag through staging. On 24 March 2026 two releases of thelitellmPython package on PyPI, 1.82.7 and 1.82.8, were malicious. LiteLLM's own incident write-up says they were live for about 40 minutes, that the official Docker proxy image was not affected because it pins its dependencies, and that users of the bad versions should rotate every secret, look forlitellm_init.pth, and pin to 1.82.6 or earlier until safe releases were confirmed. It reports releases from 1.78.0 to 1.82.6 audited, and a new CI pipeline for 1.83.0.The lesson applies to any gateway, not only this one. The gateway holds your provider keys, so it is a high-value target. Install from a pinned image or a lockfile with hashes, keep provider keys in a secrets manager, give each application a virtual key with the narrowest model list and budget, and keep the admin interface off the public internet. Plan the rotation: if you had to replace every provider key tonight, how long would it take?
Finally, decide who is paged when the gateway is down. A gateway moves a failure from many applications to one service, which is easier to watch and more damaging when it breaks. Run at least two instances behind a load balancer once real traffic depends on it.
Checked against: LiteLLM: Security Update, Suspected Supply Chain Incident, LiteLLM production best practices
LLM gateway or API gateway: what is the difference?
A general API gateway handles authentication, rate limiting and routing for any HTTP API and knows nothing about tokens. An LLM gateway adds what language model traffic needs: a single OpenAI-compatible interface across providers, spend counted in tokens and money, model-aware fallbacks, and guardrails on prompts and answers. Many teams run both: the API gateway at the edge of the network, the LLM gateway behind it.
Troubleshooting
| What you see | Likely cause | Fix |
|---|---|---|
| 401 Authentication Error when calling the gateway | The Authorization header is missing, uses the wrong key, or the master key was not passed into the container. | Send Authorization: Bearer <key>, and check the container received LITELLM_MASTER_KEY with docker run -e LITELLM_MASTER_KEY. |
| The model name in the request is rejected as not found | The request used the provider string instead of the model_name you defined, or the virtual key is not allowed that model. | Use the model_name from model_list and check the models list on the key. |
| Connection refused from the container to a local model server | Inside the container localhost is the container itself, not your machine. | Use host.docker.internal (with --add-host=host.docker.internal:host-gateway on Linux) or the host's address in api_base. |
| Budgets are not enforced, or /key/generate fails | There is no database. Budgets and virtual keys need one. | Set DATABASE_URL to a PostgreSQL connection string and restart the proxy. |
| ExceededTokenBudget error from a virtual key | The key reached its max_budget in the budget period. | Raise the budget, wait for the period to reset, or issue a new key. Make your application treat this error as a handled failure. |
| The fallback never runs | The fallback list is under a different name than the model in the request, or the failure is not one the setting covers. | Match the key in fallbacks to the model_name clients request, and trigger a failure on purpose to test it. |
| A health probe costs money or slows the service | You pointed it at /health, which sends a real request to every model. | Use /health/liveliness and /health/readiness for probes. |
Verify it worked
Next steps
- How to migrate from OpenRouter or LiteLLM: move traffic between gateways with tests, a canary and a rollback
- How to reduce LLM costs: use the gateway's usage data to cut spend in order of payoff
- How to monitor AI agents in production: trace and alert on what passes through the gateway
- LiteLLM alternatives and comparison: see how LiteLLM compares with a managed gateway
- The self-hosted stack layers explained: where a gateway sits relative to the inference server
Related guides
- How to Migrate from OpenRouter or LiteLLM (Safely): Inventory what you use, map model names, replay real requests against old and new routes, switch the base URL, then canary the traffic with a rollback ready.
- How to Self-Host an LLM with vLLM (2026 Guide): Serve an open-weight model as a private, OpenAI-compatible endpoint on your own GPU server, with memory sizing, authentication, TLS, metrics and an upgrade routine.
- How to Reduce LLM Costs: Step-by-Step Guide (2026): Measure token spend per feature first, then apply the levers in order of payoff: output caps, prompt caching, batch APIs, model routing, response caching, and a self-hosting break-even check.
- How to Monitor AI Agents in Production (2026 Guide): Give every agent run an id, record each model and tool step as a span, redact before you store, alert on loops, tool failures and cost per run, and read a weekly sample by hand.
- How to Run LLMs Locally: Ollama, LM Studio, llama.cpp: Install Ollama, LM Studio or llama.cpp, download a model that fits your memory, chat with it and call it from code through a local OpenAI-compatible endpoint.
Frequently asked questions
What is an LLM gateway?
An LLM gateway is a service that sits between your applications and your models and exposes one OpenAI-compatible API. It holds the provider keys, so applications do not, and adds routing, fallbacks, budgets, rate limits and logging across providers.
What is the LiteLLM proxy?
It is LiteLLM's self-hosted gateway. You run it in your own infrastructure, define models in a config file, and point any OpenAI-compatible client at it. Its open-source plan includes virtual keys with budgets and rate limits.
How do I set up LiteLLM with Docker?
Write a config.yaml with a model_list and a master key from an environment variable, then run the official image with the file mounted and port 4000 published. Add a PostgreSQL DATABASE_URL to enable virtual keys and budgets. Pin an image version you have tested.
Do I need a database for LiteLLM?
Not to start. A config file alone gives you a working proxy. Virtual keys and budgets need a database, and the LiteLLM documentation says budget enforcement requires one. Use PostgreSQL through the DATABASE_URL variable.
What is the difference between an LLM gateway and an API gateway?
An API gateway manages any HTTP API. An LLM gateway understands model traffic: one interface over many providers, token-based spend, model fallbacks and prompt guardrails. Teams often run an API gateway at the network edge and an LLM gateway behind it.
Is LiteLLM safe to use after the March 2026 incident?
LiteLLM says two PyPI releases, 1.82.7 and 1.82.8, were malicious on 24 March 2026 and that the official Docker proxy image was not affected. It audited earlier releases and changed its pipeline. Pin versions, install from the image, and follow its incident page if you installed from PyPI that day.
How Swfte can help
You can complete this guide with open-source software and no Swfte account. If you would rather not run the gateway yourself, Swfte Connect is a managed OpenAI-compatible gateway with routing, failover, cost tracking and budgets.
- Swfte Connect: the managed gateway: one API, routing, failover and cost analytics
- Connect self-deploy: how self-hosting the gateway is approached; scoped as an engagement
- Swfte BuildX: the model gateway as it appears in the platform
Connect is a managed service today. Private or dedicated deployment is designed for and scoped as an engagement, and the number of supported providers is <provider count - founder to fill>. If you must keep all traffic on your own network now, LiteLLM as set up above is the route to take.
Missing a step or found a command that no longer works? Tell us, or request a how-to.
Sources and last verified
Commands, versions and facts in this guide were checked against the sources below on . Tools change quickly: if something differs from what you see, trust the official documentation and let us know.
- LiteLLM gateway quickstart: Quickstart script, Docker Compose with Postgres, port 4000, admin UI, virtual keys, minimal YAML-only setup and the database requirement for budgets
- LiteLLM config.yaml documentation: model_list, model_name versus litellm_params.model, master_key from os.environ, docker run with volume mount, curl test on port 4000
- LiteLLM providers documentation: provider/model naming and the openai/ prefix with api_base for OpenAI-compatible endpoints
- LiteLLM virtual keys documentation: /key/generate with max_budget, budget_duration, models, rpm_limit, tpm_limit; DATABASE_URL; budget error messages
- LiteLLM fallbacks and reliability documentation: num_retries, request_timeout, fallbacks, allowed_fails, cooldown_time and the specialised fallback lists
- LiteLLM health endpoints: /health/liveliness, /health/readiness, and /health running real requests against every model
- LiteLLM production best practices: Sizing, LITELLM_MASTER_KEY, LITELLM_MODE, LITELLM_SALT_KEY, proxy_batch_write_at, disable_error_logs, json_logs, Redis 7.0
- LiteLLM PII masking guardrail (Presidio): guardrails block with guardrail presidio, pre_call mode, MASK and BLOCK actions
- LiteLLM: Security Update, Suspected Supply Chain Incident: Compromised versions 1.82.7 and 1.82.8 on 24 March 2026, duration, Docker proxy image not affected, recommended actions, audited safe versions
- llama.cpp server README: llama-server -m model.gguf --port 8080 and the /v1/chat/completions endpoint
Topics
- LLM gateway
- LiteLLM
- virtual keys
- fallbacks
- budgets
Machine-readable copies: this guide as markdown, index of all guides (JSON). Canonical address: https://www.swfte.com/how-to-set-up-an-llm-gateway.