Operate · Intermediate

How to set up an LLM gateway

  • Time: About 2 hours to a working, budgeted gateway; a day to production-harden it
  • Cost: LiteLLM Open Source is free to self-host. You pay your model providers and for the server and database you run it on.
  • Level: Intermediate
On this page
  1. Short answer
  2. Before you start
  3. Run it yourself or call a hosted gateway
  4. 1. Choose the first model to put behind the gateway
  5. 2. Decide how you will start the proxy
  6. 3. Write config.yaml with one model and a master key
  7. 4. Run the gateway and send a test request
  8. 5. Add Postgres and issue virtual keys with budgets
  9. 6. Set retries, timeouts and fallbacks
  10. 7. Add health checks and production settings
  11. 8. Decide what the gateway logs and redact what it should not keep
  12. 9. Harden the deployment
  13. LLM gateway or API gateway: what is the difference?
  14. Troubleshooting
  15. Verify it worked
  16. Next steps
  17. FAQ
  18. How Swfte can help
  19. Sources and last verified

Short answer

An LLM gateway is one OpenAI-compatible endpoint in front of your models, so applications hold a gateway key instead of provider keys. The quickest open-source route is the LiteLLM proxy: write a config.yaml with a model_list and a master key, run the Docker image, test with curl, add Postgres to issue virtual keys with budgets, then add fallbacks, health checks and hardening. Pin the version you run.

The steps at a glance

  1. Choose the first model to put behind the gateway
  2. Decide how you will start the proxy
  3. Write config.yaml with one model and a master key
  4. Run the gateway and send a test request
  5. Add Postgres and issue virtual keys with budgets
  6. Set retries, timeouts and fallbacks
  7. Add health checks and production settings
  8. Decide what the gateway logs and redact what it should not keep
  9. Harden the deployment

Before you start

Who this is for

  • Platform and backend engineers who want provider keys, budgets and logs in one place instead of in every application.
  • Teams running more than one model, or both hosted and self-hosted models, behind one API.
  • Anyone replacing scattered provider SDK calls with a single base URL they control.

Probably not for you if

  • Teams with one application and one provider and no budget or audit need. A gateway is another service to run, and you may not need it yet.
  • Teams that want a hosted service with no infrastructure. Compare the hosted options in the table below first.

Prerequisites

  • Docker installed, and permission to run containers on a server or your laptop.
  • At least one model to put behind the gateway: a provider API key, or an OpenAI-compatible server such as llama.cpp's llama-server.
  • A PostgreSQL database for step 5 onwards (virtual keys and budgets need a database). The LiteLLM quickstart starts one for you with Docker Compose.
  • A way to keep secrets out of files you commit: environment variables, a secrets manager, or a .env file that stays out of version control.
Time
About 2 hours to a working, budgeted gateway; a day to production-harden it
Cost
LiteLLM Open Source is free to self-host. You pay your model providers and for the server and database you run it on.
Hardware
LiteLLM's production guide recommends 1 vCPU and 4 GiB of memory per pod as a floor. A laptop is enough to follow the steps.
Skill
Comfortable with Docker, YAML and curl

Estimates are ours, not measurements, and move with your hardware, data and network.

Run it yourself or call a hosted gateway

The first decision is who carries the pager. LiteLLM is software you operate, so your prompts and keys stay on your infrastructure and so do the uptime and upgrades. A hosted gateway is a service you call, with less to run and a third party in the data path. The facts below come from the vendors' own pages as recorded on our alternatives pages on 6 October 2026. Check them before you buy.

Gateway options at a glance
OptionHow you run itPricing as publishedWorth knowing
LiteLLMOpen-source Python SDK and self-hosted proxyOpen Source: $0, self-hosted. Enterprise: sales, annualMIT licence except its enterprise directory; virtual keys with budgets and rate limits in the open-source plan
OpenRouterHosted serviceFree plan; pay-as-you-go plans with a fee on credit purchases; inference at provider list priceEU and US in-region routing is listed on the Business plan
PortkeyOpen-source gateway, hosted platform or private cloudOpen source; Developer free; Production $49 a month; Enterprise customGuardrails and semantic caching depend on the plan
HeliconeObservability proxy reached by a base-URL changeOpen-source core; free tier, then usage-based per logged requestSees only traffic that passes through the proxy
Swfte ConnectManaged gateway on the Swfte platformFree (pay as you use); Pro 5% platform fee; Scale 3 to 4%; Enterprise negotiableDedicated or private deployment is scoped as an engagement, not self-serve
  1. Step 1Choose the first model to put behind the gateway

    You end up with: One model you can already call, with its address and key written down.

    Start with a model you can already reach so you can tell a gateway fault from a model fault. Two good choices: a hosted provider where you have an API key, or a local OpenAI-compatible server. llama.cpp's llama-server is one: llama-server -m model.gguf --port 8080 serves POST /v1/chat/completions on that port.

    Write down three things: the provider or server address, the model name that backend expects, and the key (or "none" for a local server with no --api-key). You will put these in the config file in step 3. Do not put real keys in the file itself; LiteLLM reads them from environment variables, as you will see.

    If you want to test with no paid account at all, run the local server and use that. Our guide to running LLMs locally covers getting a model onto your machine.

    Optional: serve a local model on port 8080 · bash
    llama-server -m model.gguf --port 8080

    Checked against: llama.cpp server README

  2. Step 2Decide how you will start the proxy

    You end up with: A choice between the one-command quickstart and a config file you control.

    LiteLLM documents two starts. The quickstart script deploys the gateway and a Postgres database with Docker Compose, writes generated security keys to a .env file, and serves the gateway and an admin UI on http://localhost:4000. It is the fastest way to see the product. Read the script before you run it, as you should with anything you pipe into a shell.

    The second start is a YAML config file and one docker run, with no database. The documentation describes it as the minimal setup for API-only use and notes that budget enforcement needs a database. This guide uses the config file, because a file you keep in version control is repeatable and reviewable, then adds Postgres in step 5.

    If you take the quickstart route, you can still follow steps 6 to 9, and you manage models and keys in the admin UI instead of the file.

    Option A: the LiteLLM quickstart script (read it first) · bash
    curl -fsSL https://raw.githubusercontent.com/BerriAI/litellm/main/scripts/quickstart.sh | sh

    Checked against: LiteLLM gateway quickstart, LiteLLM: Security Update, Suspected Supply Chain Incident

  3. Step 3Write config.yaml with one model and a master key

    You end up with: A config file naming one model and reading its secrets from environment variables.

    The config file has three parts that matter now. model_list is the list of models clients can ask for. For each entry, model_name is the name clients use, and litellm_params.model is the provider-specific string LiteLLM passes on. general_settings.master_key is the admin credential, and the documentation shows it read from the environment as os.environ/LITELLM_MASTER_KEY.

    Model strings use a provider/model form, such as anthropic/<model> or openai/<model>. For a custom OpenAI-compatible server, such as llama.cpp, use the openai/ prefix plus an api_base. The docs show exactly this pattern for a vLLM endpoint, with api_key: none. For a hosted provider, put the key in an environment variable and refer to it as os.environ/NAME.

    Clients never see the provider key. They send the gateway key, and the gateway adds the provider key on the way out. That is most of the value of the pattern.

    config.yaml · yaml
    model_list:
      - model_name: local-llm
        litellm_params:
          model: openai/model
          api_base: http://host.docker.internal:8080/v1
          api_key: none
    
      - model_name: hosted-model
        litellm_params:
          model: <provider>/<model>
          api_key: os.environ/PROVIDER_API_KEY
    
    litellm_settings:
      drop_params: True
    
    general_settings:
      master_key: os.environ/LITELLM_MASTER_KEY
    Create a master key (it should start with sk-) and keep it out of version control · bash
    export LITELLM_MASTER_KEY="sk-$(openssl rand -hex 24)"
    export PROVIDER_API_KEY="<your provider key>"

    Checked against: LiteLLM config.yaml documentation, LiteLLM providers documentation, LiteLLM production best practices

  4. Step 4Run the gateway and send a test request

    You end up with: A completion returned through http://localhost:4000 using your master key.

    Run the official image with the config file mounted and the environment variables passed through. The documentation shows the mount form docker run -v /path/to/config.yaml:/app/config.yaml docker.litellm.ai/berriai/litellm:latest --config /app/config.yaml. Replace latest with a version tag you have tested, and publish port 4000.

    Then send a chat request with curl. The endpoint is /chat/completions on port 4000, in the OpenAI format. Because you set a master key, send it as a bearer token. The model value is your model_name from the config, not the provider string.

    Any OpenAI SDK now works by changing two values: the base URL to http://localhost:4000 and the API key to a gateway key. Application code does not otherwise change. That is how you migrate applications onto a gateway one at a time.

    Run the proxy · bash
    docker run -p 4000:4000 \
      -e LITELLM_MASTER_KEY \
      -e PROVIDER_API_KEY \
      -v "$(pwd)/config.yaml:/app/config.yaml" \
      docker.litellm.ai/berriai/litellm:<tested-version> \
      --config /app/config.yaml
    Send a request through the gateway · bash
    curl --location 'http://0.0.0.0:4000/chat/completions' \
      --header 'Content-Type: application/json' \
      --header "Authorization: Bearer $LITELLM_MASTER_KEY" \
      --data '{
        "model": "local-llm",
        "messages": [{"role": "user", "content": "Reply with the single word: ready"}]
      }'

    Checked against: LiteLLM config.yaml documentation

  5. Step 5Add Postgres and issue virtual keys with budgets

    You end up with: Per-application keys, each with its own model list, spend limit and rate limit.

    Do not give applications the master key. Create a virtual key for each application or team. LiteLLM generates them with POST /key/generate, authorised by the master key, and each key can carry a max_budget, a budget_duration, a list of allowed models, and rpm_limit and tpm_limit for requests and tokens per minute.

    Budgets and virtual keys need a database. The documentation is explicit that budgets require one, and it takes a PostgreSQL connection string in the DATABASE_URL environment variable. Start a Postgres instance (the quickstart Compose file includes one), add -e DATABASE_URL to the docker run command, and restart the proxy.

    When a key goes over its budget the gateway rejects the request with an error such as ExceededTokenBudget, which your application must handle like any other API failure. Set budgets slightly above expected use at first and tighten them once you have a month of data. Rate limits per model can also be set in model_list, for example rpm: 6 on an entry.

    Issue a virtual key (replace the values with your own) · bash
    curl 'http://0.0.0.0:4000/key/generate' \
      --header "Authorization: Bearer $LITELLM_MASTER_KEY" \
      --header 'Content-Type: application/json' \
      --data-raw '{
        "max_budget": 10,
        "budget_duration": "30d",
        "models": ["local-llm"],
        "rpm_limit": 60,
        "tpm_limit": 20000
      }'

    Checked against: LiteLLM virtual keys documentation, LiteLLM config.yaml documentation

  6. Step 6Set retries, timeouts and fallbacks

    You end up with: A failing model is retried, then replaced by the next model in your list.

    Fallbacks are the feature that makes a gateway earn its place. In litellm_settings you set num_retries, request_timeout, allowed_fails and cooldown_time, and a fallbacks list that maps a model name to the models to try next. The documentation says fallbacks run in order, and that a model which fails more than allowed_fails times in a minute is put on cooldown for cooldown_time seconds.

    There are specialised fallback lists too: context_window_fallbacks for context-length errors, content_policy_fallbacks for content policy errors and default_fallbacks for anything else. A common pattern is a small, cheap model as the primary, and a larger one as the fallback for context overflows.

    Test it. Stop the backend for the primary model, send a request, and check the answer arrives from the fallback. A fallback you have not triggered on purpose is a fallback you do not know works. Check the data protection position of every model in the chain: a fallback to a different provider changes where your data goes.

    Add to config.yaml · yaml
    litellm_settings:
      drop_params: True
      num_retries: 3
      request_timeout: 30
      fallbacks: [{"local-llm": ["hosted-model"]}]
      allowed_fails: 3
      cooldown_time: 30

    Checked against: LiteLLM fallbacks and reliability documentation

  7. Step 7Add health checks and production settings

    You end up with: Probes your orchestrator can use, and settings that suit a production deployment.

    LiteLLM exposes three health endpoints. GET /health/liveliness says the process is up and checks no dependencies. GET /health/readiness says the worker is ready for traffic, and reports the database state. GET /health runs a real test request against every configured model, so it costs a few tokens per model: do not point a once-a-second probe at it. Use the first two for container probes and the third for an occasional manual check.

    The production guide lists the settings to change: export LITELLM_MODE=PRODUCTION to stop automatic .env loading, set LITELLM_SALT_KEY to a random value to encrypt stored credentials (and never change it after deployment), write spend updates to the database every 60 seconds with proxy_batch_write_at: 60, keep error logs out of the database with disable_error_logs: True, and set json_logs: true for structured logs. For several instances it recommends Redis 7.0 or later to share rate limits and router state.

    Put TLS in front of the gateway with your reverse proxy or load balancer, and do not expose the admin UI to the internet. Keep the gateway on a network the applications can reach and the public cannot.

    Probe the gateway · bash
    curl http://localhost:4000/health/liveliness
    curl http://localhost:4000/health/readiness
    Production settings for config.yaml · yaml
    general_settings:
      master_key: os.environ/LITELLM_MASTER_KEY
      proxy_batch_write_at: 60
      disable_error_logs: True
    
    litellm_settings:
      json_logs: true

    Checked against: LiteLLM health endpoints, LiteLLM production best practices

  8. Step 8Decide what the gateway logs and redact what it should not keep

    You end up with: A written logging policy and, where needed, a PII masking guardrail in front of the model.

    A gateway sees every prompt and every answer, which is its value and its risk. Decide what you keep. At minimum you want who called, which model, token counts, cost and latency. Whether you also keep prompt and response text depends on your data protection position; keeping them makes debugging easier and makes the log a store of personal data you must protect, retain and delete properly.

    LiteLLM supports guardrails, including a Presidio-based PII masking guardrail with a pre_call mode that masks or blocks entities such as credit card numbers before the request reaches the model. The configuration is a guardrails block with a guardrail_name and litellm_params. Read the documentation page for the services Presidio needs before you rely on it, and test it on realistic text, because pattern-based masking misses things.

    Redaction in the gateway does not replace deciding what data should go to a model at all. Combine it with an allow-list of models per key and, for sensitive work, a self-hosted model. For the wider practice, see how to monitor AI agents in production.

    Guardrail block from the LiteLLM documentation (needs the Presidio services it describes) · yaml
    guardrails:
      - guardrail_name: "presidio-pii"
        litellm_params:
          guardrail: presidio
          mode: "pre_call"
          presidio_language: "en"

    Checked against: LiteLLM PII masking guardrail (Presidio)

  9. Step 9Harden the deployment

    You end up with: A pinned, minimally exposed gateway with secrets, updates and incident steps decided in advance.

    Pin versions. Run an image tag you have tested, not latest, and promote a new tag through staging. On 24 March 2026 two releases of the litellm Python package on PyPI, 1.82.7 and 1.82.8, were malicious. LiteLLM's own incident write-up says they were live for about 40 minutes, that the official Docker proxy image was not affected because it pins its dependencies, and that users of the bad versions should rotate every secret, look for litellm_init.pth, and pin to 1.82.6 or earlier until safe releases were confirmed. It reports releases from 1.78.0 to 1.82.6 audited, and a new CI pipeline for 1.83.0.

    The lesson applies to any gateway, not only this one. The gateway holds your provider keys, so it is a high-value target. Install from a pinned image or a lockfile with hashes, keep provider keys in a secrets manager, give each application a virtual key with the narrowest model list and budget, and keep the admin interface off the public internet. Plan the rotation: if you had to replace every provider key tonight, how long would it take?

    Finally, decide who is paged when the gateway is down. A gateway moves a failure from many applications to one service, which is easier to watch and more damaging when it breaks. Run at least two instances behind a load balancer once real traffic depends on it.

    Checked against: LiteLLM: Security Update, Suspected Supply Chain Incident, LiteLLM production best practices

LLM gateway or API gateway: what is the difference?

A general API gateway handles authentication, rate limiting and routing for any HTTP API and knows nothing about tokens. An LLM gateway adds what language model traffic needs: a single OpenAI-compatible interface across providers, spend counted in tokens and money, model-aware fallbacks, and guardrails on prompts and answers. Many teams run both: the API gateway at the edge of the network, the LLM gateway behind it.

Troubleshooting

What you seeLikely causeFix
401 Authentication Error when calling the gatewayThe Authorization header is missing, uses the wrong key, or the master key was not passed into the container.Send Authorization: Bearer <key>, and check the container received LITELLM_MASTER_KEY with docker run -e LITELLM_MASTER_KEY.
The model name in the request is rejected as not foundThe request used the provider string instead of the model_name you defined, or the virtual key is not allowed that model.Use the model_name from model_list and check the models list on the key.
Connection refused from the container to a local model serverInside the container localhost is the container itself, not your machine.Use host.docker.internal (with --add-host=host.docker.internal:host-gateway on Linux) or the host's address in api_base.
Budgets are not enforced, or /key/generate failsThere is no database. Budgets and virtual keys need one.Set DATABASE_URL to a PostgreSQL connection string and restart the proxy.
ExceededTokenBudget error from a virtual keyThe key reached its max_budget in the budget period.Raise the budget, wait for the period to reset, or issue a new key. Make your application treat this error as a handled failure.
The fallback never runsThe fallback list is under a different name than the model in the request, or the failure is not one the setting covers.Match the key in fallbacks to the model_name clients request, and trigger a failure on purpose to test it.
A health probe costs money or slows the serviceYou pointed it at /health, which sends a real request to every model.Use /health/liveliness and /health/readiness for probes.

Verify it worked

Next steps

Related guides

Frequently asked questions

What is an LLM gateway?

An LLM gateway is a service that sits between your applications and your models and exposes one OpenAI-compatible API. It holds the provider keys, so applications do not, and adds routing, fallbacks, budgets, rate limits and logging across providers.

What is the LiteLLM proxy?

It is LiteLLM's self-hosted gateway. You run it in your own infrastructure, define models in a config file, and point any OpenAI-compatible client at it. Its open-source plan includes virtual keys with budgets and rate limits.

How do I set up LiteLLM with Docker?

Write a config.yaml with a model_list and a master key from an environment variable, then run the official image with the file mounted and port 4000 published. Add a PostgreSQL DATABASE_URL to enable virtual keys and budgets. Pin an image version you have tested.

Do I need a database for LiteLLM?

Not to start. A config file alone gives you a working proxy. Virtual keys and budgets need a database, and the LiteLLM documentation says budget enforcement requires one. Use PostgreSQL through the DATABASE_URL variable.

What is the difference between an LLM gateway and an API gateway?

An API gateway manages any HTTP API. An LLM gateway understands model traffic: one interface over many providers, token-based spend, model fallbacks and prompt guardrails. Teams often run an API gateway at the network edge and an LLM gateway behind it.

Is LiteLLM safe to use after the March 2026 incident?

LiteLLM says two PyPI releases, 1.82.7 and 1.82.8, were malicious on 24 March 2026 and that the official Docker proxy image was not affected. It audited earlier releases and changed its pipeline. Pin versions, install from the image, and follow its incident page if you installed from PyPI that day.

How Swfte can help

You can complete this guide with open-source software and no Swfte account. If you would rather not run the gateway yourself, Swfte Connect is a managed OpenAI-compatible gateway with routing, failover, cost tracking and budgets.

  • Swfte Connect: the managed gateway: one API, routing, failover and cost analytics
  • Connect self-deploy: how self-hosting the gateway is approached; scoped as an engagement
  • Swfte BuildX: the model gateway as it appears in the platform

Connect is a managed service today. Private or dedicated deployment is designed for and scoped as an engagement, and the number of supported providers is <provider count - founder to fill>. If you must keep all traffic on your own network now, LiteLLM as set up above is the route to take.

Missing a step or found a command that no longer works? Tell us, or request a how-to.

Sources and last verified

Commands, versions and facts in this guide were checked against the sources below on . Tools change quickly: if something differs from what you see, trust the official documentation and let us know.

  1. LiteLLM gateway quickstart: Quickstart script, Docker Compose with Postgres, port 4000, admin UI, virtual keys, minimal YAML-only setup and the database requirement for budgets
  2. LiteLLM config.yaml documentation: model_list, model_name versus litellm_params.model, master_key from os.environ, docker run with volume mount, curl test on port 4000
  3. LiteLLM providers documentation: provider/model naming and the openai/ prefix with api_base for OpenAI-compatible endpoints
  4. LiteLLM virtual keys documentation: /key/generate with max_budget, budget_duration, models, rpm_limit, tpm_limit; DATABASE_URL; budget error messages
  5. LiteLLM fallbacks and reliability documentation: num_retries, request_timeout, fallbacks, allowed_fails, cooldown_time and the specialised fallback lists
  6. LiteLLM health endpoints: /health/liveliness, /health/readiness, and /health running real requests against every model
  7. LiteLLM production best practices: Sizing, LITELLM_MASTER_KEY, LITELLM_MODE, LITELLM_SALT_KEY, proxy_batch_write_at, disable_error_logs, json_logs, Redis 7.0
  8. LiteLLM PII masking guardrail (Presidio): guardrails block with guardrail presidio, pre_call mode, MASK and BLOCK actions
  9. LiteLLM: Security Update, Suspected Supply Chain Incident: Compromised versions 1.82.7 and 1.82.8 on 24 March 2026, duration, Docker proxy image not affected, recommended actions, audited safe versions
  10. llama.cpp server README: llama-server -m model.gguf --port 8080 and the /v1/chat/completions endpoint

Topics

  • LLM gateway
  • LiteLLM
  • virtual keys
  • fallbacks
  • budgets

Machine-readable copies: this guide as markdown, index of all guides (JSON). Canonical address: https://www.swfte.com/how-to-set-up-an-llm-gateway.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.