# How to set up an LLM gateway

Canonical: https://www.swfte.com/how-to-set-up-an-llm-gateway
Last verified: 2026-10-06
Difficulty: Intermediate
Time: About 2 hours to a working, budgeted gateway; a day to production-harden it
Cost: LiteLLM Open Source is free to self-host. You pay your model providers and for the server and database you run it on.
Hardware: LiteLLM's production guide recommends 1 vCPU and 4 GiB of memory per pod as a floor. A laptop is enough to follow the steps.

## Short answer

An LLM gateway is one OpenAI-compatible endpoint in front of your models, so applications hold a gateway key instead of provider keys. The quickest open-source route is the LiteLLM proxy: write a config.yaml with a model_list and a master key, run the Docker image, test with curl, add Postgres to issue virtual keys with budgets, then add fallbacks, health checks and hardening. Pin the version you run.

## Who this is for

- Platform and backend engineers who want provider keys, budgets and logs in one place instead of in every application.
- Teams running more than one model, or both hosted and self-hosted models, behind one API.
- Anyone replacing scattered provider SDK calls with a single base URL they control.

Not for:
- Teams with one application and one provider and no budget or audit need. A gateway is another service to run, and you may not need it yet.
- Teams that want a hosted service with no infrastructure. Compare the hosted options in the table below first.

## Prerequisites

- Docker installed, and permission to run containers on a server or your laptop.
- At least one model to put behind the gateway: a provider API key, or an OpenAI-compatible server such as llama.cpp's `llama-server`.
- A PostgreSQL database for step 5 onwards (virtual keys and budgets need a database). The LiteLLM quickstart starts one for you with Docker Compose.
- A way to keep secrets out of files you commit: environment variables, a secrets manager, or a `.env` file that stays out of version control.

## Run it yourself or call a hosted gateway

The first decision is who carries the pager. LiteLLM is software you operate, so your prompts and keys stay on your infrastructure and so do the uptime and upgrades. A hosted gateway is a service you call, with less to run and a third party in the data path. The facts below come from the vendors' own pages as recorded on our alternatives pages on 6 October 2026. Check them before you buy.

**Gateway options at a glance**

| Option | How you run it | Pricing as published | Worth knowing |
| --- | --- | --- | --- |
| [LiteLLM](https://www.swfte.com/alternatives/litellm) | Open-source Python SDK and self-hosted proxy | Open Source: $0, self-hosted. Enterprise: sales, annual | MIT licence except its enterprise directory; virtual keys with budgets and rate limits in the open-source plan |
| [OpenRouter](https://www.swfte.com/alternatives/openrouter) | Hosted service | Free plan; pay-as-you-go plans with a fee on credit purchases; inference at provider list price | EU and US in-region routing is listed on the Business plan |
| [Portkey](https://www.swfte.com/alternatives/portkey) | Open-source gateway, hosted platform or private cloud | Open source; Developer free; Production $49 a month; Enterprise custom | Guardrails and semantic caching depend on the plan |
| [Helicone](https://www.swfte.com/alternatives/helicone) | Observability proxy reached by a base-URL change | Open-source core; free tier, then usage-based per logged request | Sees only traffic that passes through the proxy |
| [Swfte Connect](https://www.swfte.com/products/connect) | Managed gateway on the Swfte platform | Free (pay as you use); Pro 5% platform fee; Scale 3 to 4%; Enterprise negotiable | Dedicated or private deployment is scoped as an engagement, not self-serve |

## Steps

### Step 1: Choose the first model to put behind the gateway

Outcome: One model you can already call, with its address and key written down.

Start with a model you can already reach so you can tell a gateway fault from a model fault. Two good choices: a hosted provider where you have an API key, or a local OpenAI-compatible server. llama.cpp's `llama-server` is one: `llama-server -m model.gguf --port 8080` serves `POST /v1/chat/completions` on that port.

Write down three things: the provider or server address, the model name that backend expects, and the key (or "none" for a local server with no `--api-key`). You will put these in the config file in step 3. Do not put real keys in the file itself; LiteLLM reads them from environment variables, as you will see.

If you want to test with no paid account at all, run the local server and use that. Our guide to [running LLMs locally](https://www.swfte.com/how-to-run-llms-locally) covers getting a model onto your machine.

Optional: serve a local model on port 8080:

```bash
llama-server -m model.gguf --port 8080
```

### Step 2: Decide how you will start the proxy

Outcome: A choice between the one-command quickstart and a config file you control.

LiteLLM documents two starts. The quickstart script deploys the gateway and a Postgres database with Docker Compose, writes generated security keys to a `.env` file, and serves the gateway and an admin UI on `http://localhost:4000`. It is the fastest way to see the product. Read the script before you run it, as you should with anything you pipe into a shell.

The second start is a YAML config file and one `docker run`, with no database. The documentation describes it as the minimal setup for API-only use and notes that budget enforcement needs a database. This guide uses the config file, because a file you keep in version control is repeatable and reviewable, then adds Postgres in step 5.

If you take the quickstart route, you can still follow steps 6 to 9, and you manage models and keys in the admin UI instead of the file.

Option A: the LiteLLM quickstart script (read it first):

```bash
curl -fsSL https://raw.githubusercontent.com/BerriAI/litellm/main/scripts/quickstart.sh | sh
```

> WARNING: Pin the version of the gateway you run. In March 2026 two PyPI releases of litellm were malicious; the vendor says the official Docker image was not affected, but anyone who installed from PyPI during the window was. See step 9.

### Step 3: Write config.yaml with one model and a master key

Outcome: A config file naming one model and reading its secrets from environment variables.

The config file has three parts that matter now. `model_list` is the list of models clients can ask for. For each entry, `model_name` is the name clients use, and `litellm_params.model` is the provider-specific string LiteLLM passes on. `general_settings.master_key` is the admin credential, and the documentation shows it read from the environment as `os.environ/LITELLM_MASTER_KEY`.

Model strings use a `provider/model` form, such as `anthropic/<model>` or `openai/<model>`. For a custom OpenAI-compatible server, such as llama.cpp, use the `openai/` prefix plus an `api_base`. The docs show exactly this pattern for a vLLM endpoint, with `api_key: none`. For a hosted provider, put the key in an environment variable and refer to it as `os.environ/NAME`.

Clients never see the provider key. They send the gateway key, and the gateway adds the provider key on the way out. That is most of the value of the pattern.

config.yaml:

```yaml
model_list:
  - model_name: local-llm
    litellm_params:
      model: openai/model
      api_base: http://host.docker.internal:8080/v1
      api_key: none

  - model_name: hosted-model
    litellm_params:
      model: <provider>/<model>
      api_key: os.environ/PROVIDER_API_KEY

litellm_settings:
  drop_params: True

general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
```

Create a master key (it should start with sk-) and keep it out of version control:

```bash
export LITELLM_MASTER_KEY="sk-$(openssl rand -hex 24)"
export PROVIDER_API_KEY="<your provider key>"
```

> NOTE: `host.docker.internal` lets a container on Docker Desktop reach a server on the host. On Linux, add `--add-host=host.docker.internal:host-gateway` to the `docker run` command in the next step, or use the host's address. Remove the `hosted-model` entry if you only have the local model.

### Step 4: Run the gateway and send a test request

Outcome: A completion returned through http://localhost:4000 using your master key.

Run the official image with the config file mounted and the environment variables passed through. The documentation shows the mount form `docker run -v /path/to/config.yaml:/app/config.yaml docker.litellm.ai/berriai/litellm:latest --config /app/config.yaml`. Replace `latest` with a version tag you have tested, and publish port 4000.

Then send a chat request with curl. The endpoint is `/chat/completions` on port 4000, in the OpenAI format. Because you set a master key, send it as a bearer token. The `model` value is your `model_name` from the config, not the provider string.

Any OpenAI SDK now works by changing two values: the base URL to `http://localhost:4000` and the API key to a gateway key. Application code does not otherwise change. That is how you migrate applications onto a gateway one at a time.

Run the proxy:

```bash
docker run -p 4000:4000 \
  -e LITELLM_MASTER_KEY \
  -e PROVIDER_API_KEY \
  -v "$(pwd)/config.yaml:/app/config.yaml" \
  docker.litellm.ai/berriai/litellm:<tested-version> \
  --config /app/config.yaml
```

Send a request through the gateway:

```bash
curl --location 'http://0.0.0.0:4000/chat/completions' \
  --header 'Content-Type: application/json' \
  --header "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --data '{
    "model": "local-llm",
    "messages": [{"role": "user", "content": "Reply with the single word: ready"}]
  }'
```

### Step 5: Add Postgres and issue virtual keys with budgets

Outcome: Per-application keys, each with its own model list, spend limit and rate limit.

Do not give applications the master key. Create a virtual key for each application or team. LiteLLM generates them with `POST /key/generate`, authorised by the master key, and each key can carry a `max_budget`, a `budget_duration`, a list of allowed `models`, and `rpm_limit` and `tpm_limit` for requests and tokens per minute.

Budgets and virtual keys need a database. The documentation is explicit that budgets require one, and it takes a PostgreSQL connection string in the `DATABASE_URL` environment variable. Start a Postgres instance (the quickstart Compose file includes one), add `-e DATABASE_URL` to the `docker run` command, and restart the proxy.

When a key goes over its budget the gateway rejects the request with an error such as `ExceededTokenBudget`, which your application must handle like any other API failure. Set budgets slightly above expected use at first and tighten them once you have a month of data. Rate limits per model can also be set in `model_list`, for example `rpm: 6` on an entry.

Issue a virtual key (replace the values with your own):

```bash
curl 'http://0.0.0.0:4000/key/generate' \
  --header "Authorization: Bearer $LITELLM_MASTER_KEY" \
  --header 'Content-Type: application/json' \
  --data-raw '{
    "max_budget": 10,
    "budget_duration": "30d",
    "models": ["local-llm"],
    "rpm_limit": 60,
    "tpm_limit": 20000
  }'
```

> TIP: Record which team owns each key, and rotate the master key and virtual keys on a schedule. The gateway is only as private as the keys you hand out.

### Step 6: Set retries, timeouts and fallbacks

Outcome: A failing model is retried, then replaced by the next model in your list.

Fallbacks are the feature that makes a gateway earn its place. In `litellm_settings` you set `num_retries`, `request_timeout`, `allowed_fails` and `cooldown_time`, and a `fallbacks` list that maps a model name to the models to try next. The documentation says fallbacks run in order, and that a model which fails more than `allowed_fails` times in a minute is put on cooldown for `cooldown_time` seconds.

There are specialised fallback lists too: `context_window_fallbacks` for context-length errors, `content_policy_fallbacks` for content policy errors and `default_fallbacks` for anything else. A common pattern is a small, cheap model as the primary, and a larger one as the fallback for context overflows.

Test it. Stop the backend for the primary model, send a request, and check the answer arrives from the fallback. A fallback you have not triggered on purpose is a fallback you do not know works. Check the data protection position of every model in the chain: a fallback to a different provider changes where your data goes.

Add to config.yaml:

```yaml
litellm_settings:
  drop_params: True
  num_retries: 3
  request_timeout: 30
  fallbacks: [{"local-llm": ["hosted-model"]}]
  allowed_fails: 3
  cooldown_time: 30
```

### Step 7: Add health checks and production settings

Outcome: Probes your orchestrator can use, and settings that suit a production deployment.

LiteLLM exposes three health endpoints. `GET /health/liveliness` says the process is up and checks no dependencies. `GET /health/readiness` says the worker is ready for traffic, and reports the database state. `GET /health` runs a real test request against every configured model, so it costs a few tokens per model: do not point a once-a-second probe at it. Use the first two for container probes and the third for an occasional manual check.

The production guide lists the settings to change: export `LITELLM_MODE=PRODUCTION` to stop automatic `.env` loading, set `LITELLM_SALT_KEY` to a random value to encrypt stored credentials (and never change it after deployment), write spend updates to the database every 60 seconds with `proxy_batch_write_at: 60`, keep error logs out of the database with `disable_error_logs: True`, and set `json_logs: true` for structured logs. For several instances it recommends Redis 7.0 or later to share rate limits and router state.

Put TLS in front of the gateway with your reverse proxy or load balancer, and do not expose the admin UI to the internet. Keep the gateway on a network the applications can reach and the public cannot.

Probe the gateway:

```bash
curl http://localhost:4000/health/liveliness
curl http://localhost:4000/health/readiness
```

Production settings for config.yaml:

```yaml
general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
  proxy_batch_write_at: 60
  disable_error_logs: True

litellm_settings:
  json_logs: true
```

### Step 8: Decide what the gateway logs and redact what it should not keep

Outcome: A written logging policy and, where needed, a PII masking guardrail in front of the model.

A gateway sees every prompt and every answer, which is its value and its risk. Decide what you keep. At minimum you want who called, which model, token counts, cost and latency. Whether you also keep prompt and response text depends on your data protection position; keeping them makes debugging easier and makes the log a store of personal data you must protect, retain and delete properly.

LiteLLM supports guardrails, including a Presidio-based PII masking guardrail with a `pre_call` mode that masks or blocks entities such as credit card numbers before the request reaches the model. The configuration is a `guardrails` block with a `guardrail_name` and `litellm_params`. Read the documentation page for the services Presidio needs before you rely on it, and test it on realistic text, because pattern-based masking misses things.

Redaction in the gateway does not replace deciding what data should go to a model at all. Combine it with an allow-list of models per key and, for sensitive work, a self-hosted model. For the wider practice, see [how to monitor AI agents in production](https://www.swfte.com/how-to-monitor-ai-agents-in-production).

Guardrail block from the LiteLLM documentation (needs the Presidio services it describes):

```yaml
guardrails:
  - guardrail_name: "presidio-pii"
    litellm_params:
      guardrail: presidio
      mode: "pre_call"
      presidio_language: "en"
```

### Step 9: Harden the deployment

Outcome: A pinned, minimally exposed gateway with secrets, updates and incident steps decided in advance.

Pin versions. Run an image tag you have tested, not `latest`, and promote a new tag through staging. On 24 March 2026 two releases of the `litellm` Python package on PyPI, 1.82.7 and 1.82.8, were malicious. LiteLLM's own incident write-up says they were live for about 40 minutes, that the official Docker proxy image was not affected because it pins its dependencies, and that users of the bad versions should rotate every secret, look for `litellm_init.pth`, and pin to 1.82.6 or earlier until safe releases were confirmed. It reports releases from 1.78.0 to 1.82.6 audited, and a new CI pipeline for 1.83.0.

The lesson applies to any gateway, not only this one. The gateway holds your provider keys, so it is a high-value target. Install from a pinned image or a lockfile with hashes, keep provider keys in a secrets manager, give each application a virtual key with the narrowest model list and budget, and keep the admin interface off the public internet. Plan the rotation: if you had to replace every provider key tonight, how long would it take?

Finally, decide who is paged when the gateway is down. A gateway moves a failure from many applications to one service, which is easier to watch and more damaging when it breaks. Run at least two instances behind a load balancer once real traffic depends on it.

> WARNING: If you ever installed litellm from PyPI on 24 March 2026, follow the vendor's incident page: rotate secrets, search for `litellm_init.pth` in site-packages, and pin to a safe version.

## LLM gateway or API gateway: what is the difference?

A general API gateway handles authentication, rate limiting and routing for any HTTP API and knows nothing about tokens. An LLM gateway adds what language model traffic needs: a single OpenAI-compatible interface across providers, spend counted in tokens and money, model-aware fallbacks, and guardrails on prompts and answers. Many teams run both: the API gateway at the edge of the network, the LLM gateway behind it.

## Troubleshooting

| Symptom | Likely cause | Fix |
| --- | --- | --- |
| 401 Authentication Error when calling the gateway | The Authorization header is missing, uses the wrong key, or the master key was not passed into the container. | Send `Authorization: Bearer <key>`, and check the container received `LITELLM_MASTER_KEY` with `docker run -e LITELLM_MASTER_KEY`. |
| The model name in the request is rejected as not found | The request used the provider string instead of the `model_name` you defined, or the virtual key is not allowed that model. | Use the `model_name` from `model_list` and check the `models` list on the key. |
| Connection refused from the container to a local model server | Inside the container `localhost` is the container itself, not your machine. | Use `host.docker.internal` (with `--add-host=host.docker.internal:host-gateway` on Linux) or the host's address in `api_base`. |
| Budgets are not enforced, or /key/generate fails | There is no database. Budgets and virtual keys need one. | Set `DATABASE_URL` to a PostgreSQL connection string and restart the proxy. |
| ExceededTokenBudget error from a virtual key | The key reached its `max_budget` in the budget period. | Raise the budget, wait for the period to reset, or issue a new key. Make your application treat this error as a handled failure. |
| The fallback never runs | The fallback list is under a different name than the model in the request, or the failure is not one the setting covers. | Match the key in `fallbacks` to the `model_name` clients request, and trigger a failure on purpose to test it. |
| A health probe costs money or slows the service | You pointed it at `/health`, which sends a real request to every model. | Use `/health/liveliness` and `/health/readiness` for probes. |

## Verify it worked

- [ ] A request through the gateway with a virtual key returns an answer from the model you expect.
- [ ] A request with no key, or the wrong key, is rejected.
- [ ] A virtual key over its budget is rejected with a budget error.
- [ ] Stopping the primary backend makes the request succeed through the fallback model.
- [ ] `/health/liveliness` and `/health/readiness` return healthy, and your orchestrator uses them.
- [ ] No application holds a provider key, and the image tag you run is pinned and recorded.

## Next steps

- [How to migrate from OpenRouter or LiteLLM](https://www.swfte.com/how-to-migrate-from-openrouter-or-litellm): move traffic between gateways with tests, a canary and a rollback
- [How to reduce LLM costs](https://www.swfte.com/how-to-reduce-llm-costs): use the gateway's usage data to cut spend in order of payoff
- [How to monitor AI agents in production](https://www.swfte.com/how-to-monitor-ai-agents-in-production): trace and alert on what passes through the gateway
- [LiteLLM alternatives and comparison](https://www.swfte.com/alternatives/litellm): see how LiteLLM compares with a managed gateway
- [The self-hosted stack layers explained](https://www.swfte.com/blog/self-hosted-llm-stack-layers-vllm-ollama-litellm): where a gateway sits relative to the inference server

## FAQ

### What is an LLM gateway?

An LLM gateway is a service that sits between your applications and your models and exposes one OpenAI-compatible API. It holds the provider keys, so applications do not, and adds routing, fallbacks, budgets, rate limits and logging across providers.

### What is the LiteLLM proxy?

It is LiteLLM's self-hosted gateway. You run it in your own infrastructure, define models in a config file, and point any OpenAI-compatible client at it. Its open-source plan includes virtual keys with budgets and rate limits.

### How do I set up LiteLLM with Docker?

Write a config.yaml with a model_list and a master key from an environment variable, then run the official image with the file mounted and port 4000 published. Add a PostgreSQL DATABASE_URL to enable virtual keys and budgets. Pin an image version you have tested.

### Do I need a database for LiteLLM?

Not to start. A config file alone gives you a working proxy. Virtual keys and budgets need a database, and the LiteLLM documentation says budget enforcement requires one. Use PostgreSQL through the DATABASE_URL variable.

### What is the difference between an LLM gateway and an API gateway?

An API gateway manages any HTTP API. An LLM gateway understands model traffic: one interface over many providers, token-based spend, model fallbacks and prompt guardrails. Teams often run an API gateway at the network edge and an LLM gateway behind it.

### Is LiteLLM safe to use after the March 2026 incident?

LiteLLM says two PyPI releases, 1.82.7 and 1.82.8, were malicious on 24 March 2026 and that the official Docker proxy image was not affected. It audited earlier releases and changed its pipeline. Pin versions, install from the image, and follow its incident page if you installed from PyPI that day.

## How Swfte can help

You can complete this guide with open-source software and no Swfte account. If you would rather not run the gateway yourself, Swfte Connect is a managed OpenAI-compatible gateway with routing, failover, cost tracking and budgets.

- [Swfte Connect](https://www.swfte.com/products/connect): the managed gateway: one API, routing, failover and cost analytics
- [Connect self-deploy](https://www.swfte.com/products/connect/self-deploy): how self-hosting the gateway is approached; scoped as an engagement
- [Swfte BuildX](https://www.swfte.com/products/buildx): the model gateway as it appears in the platform

Connect is a managed service today. Private or dedicated deployment is designed for and scoped as an engagement, and the number of supported providers is <provider count - founder to fill>. If you must keep all traffic on your own network now, LiteLLM as set up above is the route to take.

## Sources

- [LiteLLM gateway quickstart](https://docs.litellm.ai/docs/proxy/docker_quick_start): Quickstart script, Docker Compose with Postgres, port 4000, admin UI, virtual keys, minimal YAML-only setup and the database requirement for budgets
- [LiteLLM config.yaml documentation](https://docs.litellm.ai/docs/proxy/configs): model_list, model_name versus litellm_params.model, master_key from os.environ, docker run with volume mount, curl test on port 4000
- [LiteLLM providers documentation](https://docs.litellm.ai/docs/providers): provider/model naming and the openai/ prefix with api_base for OpenAI-compatible endpoints
- [LiteLLM virtual keys documentation](https://docs.litellm.ai/docs/proxy/users): /key/generate with max_budget, budget_duration, models, rpm_limit, tpm_limit; DATABASE_URL; budget error messages
- [LiteLLM fallbacks and reliability documentation](https://docs.litellm.ai/docs/proxy/reliability): num_retries, request_timeout, fallbacks, allowed_fails, cooldown_time and the specialised fallback lists
- [LiteLLM health endpoints](https://docs.litellm.ai/docs/proxy/health): /health/liveliness, /health/readiness, and /health running real requests against every model
- [LiteLLM production best practices](https://docs.litellm.ai/docs/proxy/prod): Sizing, LITELLM_MASTER_KEY, LITELLM_MODE, LITELLM_SALT_KEY, proxy_batch_write_at, disable_error_logs, json_logs, Redis 7.0
- [LiteLLM PII masking guardrail (Presidio)](https://docs.litellm.ai/docs/proxy/guardrails/pii_masking_v2): guardrails block with guardrail presidio, pre_call mode, MASK and BLOCK actions
- [LiteLLM: Security Update, Suspected Supply Chain Incident](https://docs.litellm.ai/blog/security-update-march-2026): Compromised versions 1.82.7 and 1.82.8 on 24 March 2026, duration, Docker proxy image not affected, recommended actions, audited safe versions
- [llama.cpp server README](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md): llama-server -m model.gguf --port 8080 and the /v1/chat/completions endpoint

Last verified against these sources on 2026-10-06.
