# How to migrate from OpenRouter or LiteLLM

Canonical: https://www.swfte.com/how-to-migrate-from-openrouter-or-litellm
Last verified: 2026-10-06
Difficulty: Intermediate
Time: About 1 to 2 days of work for one application, plus a canary period of a few days
Cost: The cost of replaying a few hundred requests on both routes, plus any overlap in provider spend during the canary.

## Short answer

Changing a base URL takes two lines, and the risk is in everything else: model names, tool calls, structured output, streaming, usage fields, routing rules and cost. Inventory what you use, save a replay set of real requests, compare old and new routes on it, switch the client behind a percentage flag, watch cost and errors, and keep the old route until you are sure. Do not migrate if your reason is only habit.

## Who this is for

- Engineers whose application calls OpenRouter or a LiteLLM proxy and who have decided to move to direct provider APIs, another gateway, or Swfte Connect.
- Platform teams consolidating several gateways, or moving from a hosted gateway to one they run.
- Anyone asked "can we just change the URL?" who wants the evidence behind the answer.

Not for:
- Teams building a first gateway. Start with [how to set up an LLM gateway](https://www.swfte.com/how-to-set-up-an-llm-gateway).
- Teams whose current setup works and whose only reason to move is novelty. Read the next section before you spend a week on this.

## Prerequisites

- Access to the code and configuration that call the current gateway, and the ability to deploy a change behind a flag.
- Keys or accounts for the destination, and permission to send test traffic that costs money.
- A way to read your current usage: request logs, the gateway's activity export, or the provider dashboard.
- Python 3 and the `openai` package for the parity script (`pip install openai`).
- Agreement on who approves the cut-over and who can roll it back.

## When not to migrate

Most migration pages on this topic are written by the gateway that wants your traffic, and they describe the move as two lines of code. Two lines are the cheap part. Before you start, check whether your reason holds up.

- **You use several providers through one key and value that.** Moving to one provider's API gives up the single invoice and the fallback across providers. OpenRouter's breadth is its point.
- **The saving is small against the engineering time.** Our [OpenRouter against direct Anthropic comparison](https://www.swfte.com/compare/openrouter-vs-anthropic-direct) frames the fee as the price of flexibility. Do your own arithmetic with your volume.
- **You would be trading a service you do not operate for software you must.** LiteLLM gives you control, and you carry upgrades, security and uptime. See [how to set up an LLM gateway](https://www.swfte.com/how-to-set-up-an-llm-gateway) for what that involves, including the March 2026 supply-chain incident.
- **You rely on a feature the destination lacks.** For example, request-level provider routing, zero-data-retention routing, or a schema guarantee.
- **Nothing is wrong today.** A working route is an asset. Migrate for a reason you can name: cost, control, data location, a feature, or consolidation.

## Steps

### Step 1: Inventory what you actually use

Outcome: A list of every model, feature and gateway-specific field your applications rely on, with a rough volume for each.

Start from facts, not from memory. Search the code for the gateway's base URL and for any code that adds gateway-specific fields to requests. Then pull a month of usage from the gateway or provider and group it by model, by application and by feature: streaming, tool calls, JSON output, image input, long prompts.

Write the result in a table with one row per model and application. The columns that matter most are request volume, whether the call streams, whether it uses tools, whether it relies on JSON output, and whether it uses any routing preference. Volume tells you where a saving or a mistake is large. The feature columns tell you where the destination might behave differently.

Also list the non-code dependencies: dashboards that read the gateway's activity logs, budget alerts, and any contract or data processing terms. A migration that forgets the finance dashboard fails in month two.

Find calls to the current gateway:

```bash
grep -rn "openrouter.ai" --include="*.py" --include="*.ts" --include="*.js" --include="*.yaml" --include="*.json" .
grep -rn "OPENROUTER_API_KEY\|LITELLM" --include="*.py" --include="*.ts" --include="*.js" --include="*.yaml" --include="*.env.example" .
```

**Inventory template**

| Application | Model | Requests per month | Streams? | Tools? | JSON output? | Routing preference? |
| --- | --- | --- | --- | --- | --- | --- |
| support-bot | record the model id as sent today | from usage export | yes or no | yes or no | yes or no | provider order, fallbacks, zdr |

### Step 2: Map model names and request fields

Outcome: For every model in the inventory, the exact identifier the destination expects, and a note for each field that does not carry over.

Model names are the first thing to break. OpenRouter model ids are of the form `provider/model`, for example the `~openai/gpt-sol-latest` style alias its quickstart uses, or a specific identifier from its model catalogue. LiteLLM also names models `provider/model`, but there that string goes in `litellm_params.model`, while clients use the `model_name` alias you defined. Direct provider APIs use their own native names. The Swfte gateway uses `provider:model`, and the model directory shows the id for each model.

Then map the request fields. OpenRouter accepts a `provider` object (`order`, `allow_fallbacks`, `only`, `ignore`, `sort`, `zdr`, `data_collection`, `require_parameters`) and a `models` array for fallbacks. These are OpenRouter features. They do not exist on a direct provider API, and on a LiteLLM or Swfte route you re-express them as gateway configuration. Make a row for each one you use.

Optional OpenRouter headers (`HTTP-Referer` and `X-OpenRouter-Title`) only attribute your app on its leaderboards. Remove them when you leave unless the destination asks for something similar.

**Identifier forms by destination**

| Destination | Model identifier form | Where gateway features live |
| --- | --- | --- |
| OpenRouter | `provider/model` or an alias such as `~provider/name-latest` | `provider` object and `models` array in the request |
| LiteLLM proxy | Client uses your `model_name`; config maps it to `provider/model` | `fallbacks`, `num_retries` and keys in `config.yaml` |
| Direct provider (OpenAI-compatible endpoint) | The provider's own model name | Nothing: you build routing and fallback yourself |
| Swfte Connect | `provider:model` | Routing and failover settings in the platform |

### Step 3: Save a replay set of real requests

Outcome: A file of 100 to 300 real requests, with personal data removed, that you can send to any route.

The only honest answer to "does it behave the same?" is to send the same requests down both routes. Export requests from your logs, sample across applications and features, and make sure the sample includes your hard cases: the longest prompts, requests with tools, requests that expect JSON, requests in each language you serve, and a handful that currently fail.

Remove or replace personal data before you save the file, and treat the file as sensitive anyway. Keep one JSON object per line with the messages, the options you send (tools, response format, temperature), and a label for the feature it comes from. Do not store the old route's answers as the truth: they are one sample of a non-deterministic system. You compare behaviour, not exact text.

If you do not log request bodies today, you will have to write representative requests by hand. That is slower and less honest, and it is a good reason to add logging with redaction as part of this project.

replay.jsonl (one request per line; write real ones):

```json
{"feature": "support-reply", "model": "<model alias>", "messages": [{"role": "user", "content": "Reply in two sentences: a customer asks for a refund after 45 days."}], "stream": false}
{"feature": "extract-json", "model": "<model alias>", "messages": [{"role": "user", "content": "Return JSON with keys name and date for: Ana booked on 3 May."}], "stream": false}
```

### Step 4: Compare old and new routes on the replay set

Outcome: A table of differences in success, finish reason, tool calls, JSON validity, token usage and latency between the two routes.

Write a small script that sends every replay request to the old route and the new route and records what you care about: whether the call succeeded, the HTTP error if not, the `finish_reason`, whether tool calls came back, whether JSON output parses, the token counts the route reports, and the time taken. The script below does that with the OpenAI SDK pointed at two base URLs. Run it with identical settings and at temperature 0 where the destination allows it.

Read the differences, not the average. The failures that hurt are specific: a tool call that comes back malformed, a JSON answer wrapped in prose, a request rejected because the destination does not support a parameter, a missing usage field that breaks your cost tracking. Fix each one or decide it is acceptable, and write the decision down.

Expect some differences in wording. They are not failures unless a downstream rule depends on the exact text. If it does, that rule is fragile and now is a good time to fix it.

parity.py:

```python
import json
import os
import sys
import time

from openai import OpenAI

OLD = OpenAI(base_url=os.environ["OLD_BASE_URL"], api_key=os.environ["OLD_API_KEY"])
NEW = OpenAI(base_url=os.environ["NEW_BASE_URL"], api_key=os.environ["NEW_API_KEY"])
MODEL_MAP = json.loads(os.environ.get("MODEL_MAP", "{}"))  # old model id -> new model id


def call(client, model, req):
    started = time.perf_counter()
    try:
        r = client.chat.completions.create(
            model=model,
            messages=req["messages"],
            temperature=0,
            max_tokens=req.get("max_tokens", 400),
            **({"tools": req["tools"]} if "tools" in req else {}),
        )
    except Exception as exc:  # record the failure, keep going
        return {"ok": False, "error": f"{type(exc).__name__}: {exc}"[:200]}
    choice = r.choices[0]
    text = choice.message.content or ""
    return {
        "ok": True,
        "finish": choice.finish_reason,
        "tool_calls": len(choice.message.tool_calls or []),
        "json_ok": _is_json(text),
        "usage": r.usage.model_dump() if r.usage else None,
        "seconds": round(time.perf_counter() - started, 2),
    }


def _is_json(text):
    try:
        json.loads(text)
        return True
    except ValueError:
        return False


for line in open(sys.argv[1], encoding="utf-8"):
    req = json.loads(line)
    old = call(OLD, req["model"], req)
    new = call(NEW, MODEL_MAP.get(req["model"], req["model"]), req)
    diffs = [k for k in ("ok", "finish", "tool_calls", "json_ok") if old.get(k) != new.get(k)]
    print(json.dumps({"feature": req["feature"], "differs_on": diffs, "old": old, "new": new}))
```

Run it:

```bash
export OLD_BASE_URL="https://openrouter.ai/api/v1"
export OLD_API_KEY="$OPENROUTER_API_KEY"
export NEW_BASE_URL="<destination base url>"
export NEW_API_KEY="<destination key>"
export MODEL_MAP='{"<old id>": "<new id>"}'
python parity.py replay.jsonl > parity.jsonl
```

### Step 5: Change the base URL, key and model names

Outcome: The client code points at the destination behind a configuration switch, with the old route still available.

With the parity results in hand, change the client. For any OpenAI-compatible destination the change is three values: the base URL, the API key and the model identifier. Put them in configuration, not in code, so the same build can talk to either route. OpenRouter's base URL is `https://openrouter.ai/api/v1`. A LiteLLM proxy is whatever address you deployed it at, and uses your key and your `model_name` aliases. Swfte's gateway endpoint is `https://api.swfte.com/agents/v2/gateway`, with `provider:model` ids.

If your destination is a direct provider, read its compatibility notes first. Anthropic documents an OpenAI SDK compatibility layer at `https://api.anthropic.com/v1/`, and says it is mainly intended for testing and comparing models, not as a long-term production solution for most use cases. It lists differences that matter: the `strict` parameter for function calling is ignored, so tool-call JSON is not guaranteed to follow your schema; `response_format` is ignored; prompt caching is not supported through it; system and developer messages are hoisted and joined into one; `n` must be 1; `temperature` above 1 is capped at 1; and many unsupported fields are silently ignored rather than rejected. For production use of Claude directly, use Anthropic's own SDK and API.

Silent ignoring is the dangerous part. A request that used to enforce a JSON schema through the gateway may now return free text with no error. Your parity test should have caught it; if it did not, add a case for it now.

Configuration-driven client (Python):

```python
import os

from openai import OpenAI

client = OpenAI(
    base_url=os.environ["LLM_BASE_URL"],  # was https://openrouter.ai/api/v1
    api_key=os.environ["LLM_API_KEY"],
)
MODEL = os.environ["LLM_MODEL"]  # use the destination's identifier form
```

Test one request against the Swfte gateway endpoint (from the Swfte alternatives pages):

```bash
curl https://api.swfte.com/agents/v2/gateway/chat/completions \
  -H "Authorization: Bearer $SWFTE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai:gpt-4o",
    "messages": [{"role": "user", "content": "Hello"}]
  }'
```

> WARNING: Moving from a multi-provider gateway to one provider's OpenAI-compatible endpoint removes the cross-provider fallback you had. If the provider has an outage, nothing catches the request unless you build that yourself.

### Step 6: Re-express the routing and privacy settings you relied on

Outcome: Each OpenRouter or LiteLLM routing rule has an equivalent on the destination, or a recorded decision to drop it.

List what the old route did for you without being asked. On OpenRouter, a `models` array falls back to the next model when the first fails, with the documented triggers of context-length errors, moderation flags on filtered models, rate limiting and downtime, and the request is priced at the model actually used, reported in the `model` field of the response. The `provider` object can restrict providers, set an order, prefer low price, throughput or latency, or send only to zero-data-retention endpoints. In-region routing for the EU or US is described as available on its Business and Enterprise plans.

On a LiteLLM proxy the equivalents are `fallbacks`, `num_retries`, `allowed_fails` and `cooldown_time` in the config file, and per-key model allow-lists. On other gateways the setting names differ. For every rule, write the destination's equivalent or write "dropped, because ...". Privacy rules need extra care: if you relied on `zdr` or `data_collection: deny`, confirm the destination can give you the same guarantee in writing, not only a similar setting name.

If you move to a single provider, cross-provider fallback and price-based routing are gone. If you move to a self-hosted gateway, you inherit the operation of those features; see [how to set up an LLM gateway](https://www.swfte.com/how-to-set-up-an-llm-gateway).

**Examples of settings to re-express**

| If you used this on OpenRouter | It did this | Where to look on the destination |
| --- | --- | --- |
| `models: [a, b]` | Fall back to the next model on listed errors | Gateway fallback list, or your own retry code |
| `provider.order`, `only`, `ignore` | Choose and exclude upstream providers | Gateway routing rules, or the provider you chose |
| `provider.sort` by price, throughput, latency | Prefer a cheaper or faster upstream | Gateway routing strategy; may not exist |
| `provider.zdr`, `data_collection: deny` | Restrict to zero-data-retention or non-collecting endpoints | Contract terms and settings; get it in writing |
| `eu.openrouter.ai` endpoint | In-region routing (Business and Enterprise plans) | Destination's region options and contract |

### Step 7: Test streaming, tool calls and structured output on their own

Outcome: Each of the three features has a passing test on the new route, written down as a regression test.

Streaming is where hand-written parsers break. OpenRouter documents that it periodically sends Server-Sent Events comments such as `: OPENROUTER PROCESSING` to prevent timeouts, tells clients to skip lines that start with a colon before parsing JSON, ends every Chat Completions stream with an extra chunk that carries `usage`, and, once headers are sent, reports mid-stream errors as events with an `error` object and `finish_reason: "error"`. If your parser handles those, check that the destination does not need different handling. If you used the official SDK, you are likely fine, but test.

For tools, test a call that must use a tool, a call that must not, and a call with several tools. Compare the shape of `tool_calls`, argument JSON validity and how your code handles a refusal. For structured output, check whether the destination enforces your schema or only asks for it, using a case that tempts the model to add commentary.

Turn each of these into a permanent test that runs against whichever route is live. You will need them again at the next migration.

> TIP: Include one deliberately broken request, such as an unsupported parameter or an oversized prompt, and check the error your code receives. Error shapes differ between gateways and between a gateway and a provider, and error handling is untested in most applications.

### Step 8: Match cost tracking and logging

Outcome: Spend, usage and audit data from the new route lands where your dashboards and alerts expect it.

Compare the parity run's `usage` fields. Some routes fill details such as cached tokens and reasoning tokens, and some leave them empty: Anthropic's compatibility layer documents `usage.prompt_tokens_details` and `usage.completion_tokens_details` as always empty. If your cost model reads those fields, it will under-report on the new route until you change it.

Then compare price. Hosted gateways typically pass provider list prices through and add a fee. Our [OpenRouter alternatives page](https://www.swfte.com/alternatives/openrouter) records a percentage fee on credit purchases that depends on the plan, so check the current figures on OpenRouter's pricing page against your volume before you decide the migration saves money. A direct route removes that fee but also removes the single invoice and the cross-provider fallback. Work out the monthly difference from your own inventory, with the arithmetic shown, rather than from a comparison page. Our [cost reduction guide](https://www.swfte.com/how-to-reduce-llm-costs) shows how.

Finally, logging and audit. Move the alerts, budgets and dashboards that depended on the old gateway's activity data, and decide how long you keep the old data. Switch off the old billing last, and only after the final invoice has matched your records.

### Step 9: Canary the traffic and keep a rollback

Outcome: A growing share of live traffic on the new route, with a tested one-step way back.

Send a small share of real traffic, such as 1 per cent, to the new route, and raise it in steps while you watch. Make the choice sticky per user or session so one conversation does not switch route halfway. Keep the old route configured and warm until the end. A configuration flag that flips back in one step is your rollback; test it before the canary starts, not during an incident.

Decide the numbers that stop the rollout before you start: for example error rate above the old route's by a set margin, p95 latency above a limit, a drop in your task pass rate, or spend per request above your estimate. Watch them at each step for long enough to cover a full cycle of your traffic, such as a weekday and a weekend if usage differs.

When you reach 100 per cent, leave the old route in place for a defined period, then remove it, delete unused keys, and close the old account or plan. Update the documents that name the old gateway. Record the migration, the evidence and the date, so the next person can see why the decision was made.

Sticky percentage canary (Python):

```python
import hashlib
import os


def use_new_route(user_id: str) -> bool:
    """Same user always gets the same route for a given percentage."""
    percent = int(os.environ.get("NEW_ROUTE_PERCENT", "0"))  # 0 is the rollback
    bucket = int(hashlib.sha256(user_id.encode("utf-8")).hexdigest(), 16) % 100
    return bucket < percent
```

## Reasons that do justify it

Data location and control: you need traffic to stay in a region or on your own network. Cost at volume: your spend is large enough that a fee matters, and you use one or two models. Consolidation: three teams run three gateways. A feature you need: budgets per team, guardrails, audit logs. Vendor risk: you want an exit plan. Each of these has a measurable test in the steps above. See our alternatives pages for [OpenRouter](https://www.swfte.com/alternatives/openrouter) and [LiteLLM](https://www.swfte.com/alternatives/litellm) for side-by-side facts, each dated.

## Troubleshooting

| Symptom | Likely cause | Fix |
| --- | --- | --- |
| 404 or "model not found" after switching the base URL | The model identifier is still in the old gateway's form, such as `provider/model`. | Map every id to the destination's form (step 2). On Swfte use `provider:model`; on LiteLLM use your `model_name` alias. |
| JSON responses arrive as prose, or schemas stop being enforced | The destination ignores `response_format` or the `strict` flag. Anthropic's compatibility layer documents both as ignored. | Use the destination's native structured-output feature, or validate and retry in your code. Add a regression test. |
| Streaming parser throws on the first lines of the stream | The stream contains SSE comment lines such as `: OPENROUTER PROCESSING`, or differs in how it ends and reports usage. | Skip lines that start with a colon before parsing, or use the official SDK. Test the destination's stream on its own. |
| Cost dashboards drop to zero or under-report after the switch | The usage fields your dashboard reads are empty or named differently on the new route. | Compare `usage` from the parity run and update the cost code. Compute cost from prompt and completion tokens as a fallback. |
| Tool calls fail or loop on the new route | Argument JSON is not guaranteed to match your schema, or the model behaves differently with your tool descriptions. | Validate arguments server-side, return errors to the model as tool results, and compare tool-call tests on both routes. |
| Requests succeed, but quality complaints rise after the canary grows | Different default parameters, a different model version behind a name, or a prompt tuned to the old model. | Pin model versions, compare sampling parameters, and rerun your task set. Roll back with the flag while you investigate. |

## Verify it worked

- [ ] The inventory lists every model, application and gateway-specific field, with volumes.
- [ ] The parity script ran on at least 100 real requests, and every difference has a recorded decision.
- [ ] Streaming, tool calls and structured output each have a test that passes on the new route.
- [ ] Cost and usage data from the new route reach your dashboards, and the monthly cost estimate uses your own volumes.
- [ ] The rollback flag was tested before the canary and returns all traffic to the old route in one step.
- [ ] The stop conditions for the canary were written before it started, and the old account was closed only after the last invoice matched.

## Next steps

- [How to set up an LLM gateway](https://www.swfte.com/how-to-set-up-an-llm-gateway): run the gateway yourself if that is where you are moving
- [How to reduce LLM costs](https://www.swfte.com/how-to-reduce-llm-costs): do the cost arithmetic from your own usage
- [OpenRouter alternatives](https://www.swfte.com/alternatives/openrouter): dated facts on OpenRouter and what Swfte Connect offers
- [LiteLLM alternatives](https://www.swfte.com/alternatives/litellm): dated facts on LiteLLM and how it compares
- [OpenRouter or Anthropic direct](https://www.swfte.com/compare/openrouter-vs-anthropic-direct): the cost and flexibility trade-off for Claude workloads

## FAQ

### How do I migrate from OpenRouter?

Inventory the models and features you use, map model ids and routing settings to the destination, replay real requests against both routes, change the base URL, key and model names behind a flag, move traffic in steps with a rollback ready, and match cost tracking before you close the account.

### Is migrating just a base URL change?

The code change is small, and the behaviour change is not. Model ids, tool calls, JSON output, streaming, usage fields, routing rules and privacy settings can all differ. A replay test shows which of those affect your application before users do.

### What is the OpenRouter base URL and model id format?

The base URL is https://openrouter.ai/api/v1. Model ids take the form provider/model, or an alias such as ~openai/gpt-sol-latest from its quickstart, and the catalogue lists the exact identifiers. Optional headers HTTP-Referer and X-OpenRouter-Title only attribute your app on its leaderboards.

### Can I use the OpenAI SDK with Claude directly?

Yes, through Anthropic's compatibility layer at https://api.anthropic.com/v1/. Anthropic says it is mainly for testing and comparison, and lists differences: strict function-calling is ignored, response_format is ignored, prompt caching is not supported, and system messages are joined. For production, use the native API.

### How do I replace OpenRouter provider fallbacks?

OpenRouter falls back through a models array or its provider settings. On LiteLLM use the fallbacks list with retries and cooldown in config.yaml. On a single provider you write retry and fallback logic yourself. Test it by stopping the primary route on purpose.

### Should I move from LiteLLM to a managed gateway?

Move if running the proxy costs more in attention than it saves, or you need features it lacks. Stay if control, data location and no per-request fee matter more. List what you operate today, including upgrades and key rotation, and compare that with the managed price.

## How Swfte can help

You can complete this migration to any destination without Swfte. If Swfte Connect is on your shortlist, it is a managed OpenAI-compatible gateway with routing, failover, cost tracking and budgets, and our alternatives pages set out the comparison with dates.

- [Swfte Connect](https://www.swfte.com/products/connect): the managed gateway and what it covers
- [Moving from OpenRouter to Swfte](https://www.swfte.com/alternatives/openrouter): the side-by-side facts and a short migration outline
- [Moving from LiteLLM to Swfte](https://www.swfte.com/alternatives/litellm): what changes if you stop running the proxy yourself

Connect is a managed service. SAML SSO and SCIM are in development, and dedicated or private deployment is scoped as an engagement, not self-serve. If you need either today, a self-hosted gateway or a direct route may fit better. Supported provider count: <provider count - founder to fill>.

## Sources

- [OpenRouter quickstart](https://openrouter.ai/docs/quickstart): Base URL https://openrouter.ai/api/v1, OpenAI SDK setup, HTTP-Referer and X-OpenRouter-Title headers, model id format, curl example
- [OpenRouter provider routing](https://openrouter.ai/docs/guides/routing/provider-selection): provider object fields (order, allow_fallbacks, require_parameters, data_collection, zdr, only, ignore, sort) and EU and US in-region routing on Business and Enterprise plans
- [OpenRouter model fallbacks](https://openrouter.ai/docs/guides/routing/model-fallbacks): models array fallback, triggers, and the model field reporting which model was used and priced
- [OpenRouter streaming](https://openrouter.ai/docs/api-reference/streaming): SSE comment payloads, usage chunk at the end of the stream, mid-stream error format
- [LiteLLM providers documentation](https://docs.litellm.ai/docs/providers): provider/model naming and the openai/ prefix with api_base for OpenAI-compatible endpoints
- [LiteLLM fallbacks and reliability documentation](https://docs.litellm.ai/docs/proxy/reliability): fallbacks, num_retries, allowed_fails and cooldown_time
- [Claude API: OpenAI SDK compatibility](https://platform.claude.com/docs/en/api/openai-sdk): Base URL, intended use, and the list of ignored or unsupported fields including strict, response_format, prompt caching, system message hoisting, n, temperature cap and empty usage details
- [Swfte alternatives pages (gateway endpoint and model id form)](https://www.swfte.com/alternatives/openrouter): Swfte gateway endpoint, provider:model id form and the migration outline, as published on swfte.com

Last verified against these sources on 2026-10-06.
