Operate · Intermediate

How to migrate from OpenRouter or LiteLLM

  • Time: About 1 to 2 days of work for one application, plus a canary period of a few days
  • Cost: The cost of replaying a few hundred requests on both routes, plus any overlap in provider spend during the canary.
  • Level: Intermediate
On this page
  1. Short answer
  2. Before you start
  3. When not to migrate
  4. 1. Inventory what you actually use
  5. 2. Map model names and request fields
  6. 3. Save a replay set of real requests
  7. 4. Compare old and new routes on the replay set
  8. 5. Change the base URL, key and model names
  9. 6. Re-express the routing and privacy settings you relied on
  10. 7. Test streaming, tool calls and structured output on their own
  11. 8. Match cost tracking and logging
  12. 9. Canary the traffic and keep a rollback
  13. Reasons that do justify it
  14. Troubleshooting
  15. Verify it worked
  16. Next steps
  17. FAQ
  18. How Swfte can help
  19. Sources and last verified

Short answer

Changing a base URL takes two lines, and the risk is in everything else: model names, tool calls, structured output, streaming, usage fields, routing rules and cost. Inventory what you use, save a replay set of real requests, compare old and new routes on it, switch the client behind a percentage flag, watch cost and errors, and keep the old route until you are sure. Do not migrate if your reason is only habit.

The steps at a glance

  1. Inventory what you actually use
  2. Map model names and request fields
  3. Save a replay set of real requests
  4. Compare old and new routes on the replay set
  5. Change the base URL, key and model names
  6. Re-express the routing and privacy settings you relied on
  7. Test streaming, tool calls and structured output on their own
  8. Match cost tracking and logging
  9. Canary the traffic and keep a rollback

Before you start

Who this is for

  • Engineers whose application calls OpenRouter or a LiteLLM proxy and who have decided to move to direct provider APIs, another gateway, or Swfte Connect.
  • Platform teams consolidating several gateways, or moving from a hosted gateway to one they run.
  • Anyone asked "can we just change the URL?" who wants the evidence behind the answer.

Probably not for you if

  • Teams building a first gateway. Start with how to set up an LLM gateway.
  • Teams whose current setup works and whose only reason to move is novelty. Read the next section before you spend a week on this.

Prerequisites

  • Access to the code and configuration that call the current gateway, and the ability to deploy a change behind a flag.
  • Keys or accounts for the destination, and permission to send test traffic that costs money.
  • A way to read your current usage: request logs, the gateway's activity export, or the provider dashboard.
  • Python 3 and the openai package for the parity script (pip install openai).
  • Agreement on who approves the cut-over and who can roll it back.
Time
About 1 to 2 days of work for one application, plus a canary period of a few days
Cost
The cost of replaying a few hundred requests on both routes, plus any overlap in provider spend during the canary.
Skill
Comfortable changing production API client code and reading usage logs

Estimates are ours, not measurements, and move with your hardware, data and network.

When not to migrate

Most migration pages on this topic are written by the gateway that wants your traffic, and they describe the move as two lines of code. Two lines are the cheap part. Before you start, check whether your reason holds up.

  • You use several providers through one key and value that. Moving to one provider's API gives up the single invoice and the fallback across providers. OpenRouter's breadth is its point.
  • The saving is small against the engineering time. Our OpenRouter against direct Anthropic comparison frames the fee as the price of flexibility. Do your own arithmetic with your volume.
  • You would be trading a service you do not operate for software you must. LiteLLM gives you control, and you carry upgrades, security and uptime. See how to set up an LLM gateway for what that involves, including the March 2026 supply-chain incident.
  • You rely on a feature the destination lacks. For example, request-level provider routing, zero-data-retention routing, or a schema guarantee.
  • Nothing is wrong today. A working route is an asset. Migrate for a reason you can name: cost, control, data location, a feature, or consolidation.
  1. Step 1Inventory what you actually use

    You end up with: A list of every model, feature and gateway-specific field your applications rely on, with a rough volume for each.

    Start from facts, not from memory. Search the code for the gateway's base URL and for any code that adds gateway-specific fields to requests. Then pull a month of usage from the gateway or provider and group it by model, by application and by feature: streaming, tool calls, JSON output, image input, long prompts.

    Write the result in a table with one row per model and application. The columns that matter most are request volume, whether the call streams, whether it uses tools, whether it relies on JSON output, and whether it uses any routing preference. Volume tells you where a saving or a mistake is large. The feature columns tell you where the destination might behave differently.

    Also list the non-code dependencies: dashboards that read the gateway's activity logs, budget alerts, and any contract or data processing terms. A migration that forgets the finance dashboard fails in month two.

    Find calls to the current gateway · bash
    grep -rn "openrouter.ai" --include="*.py" --include="*.ts" --include="*.js" --include="*.yaml" --include="*.json" .
    grep -rn "OPENROUTER_API_KEY\|LITELLM" --include="*.py" --include="*.ts" --include="*.js" --include="*.yaml" --include="*.env.example" .
    Inventory template
    ApplicationModelRequests per monthStreams?Tools?JSON output?Routing preference?
    support-botrecord the model id as sent todayfrom usage exportyes or noyes or noyes or noprovider order, fallbacks, zdr
  2. Step 2Map model names and request fields

    You end up with: For every model in the inventory, the exact identifier the destination expects, and a note for each field that does not carry over.

    Model names are the first thing to break. OpenRouter model ids are of the form provider/model, for example the ~openai/gpt-sol-latest style alias its quickstart uses, or a specific identifier from its model catalogue. LiteLLM also names models provider/model, but there that string goes in litellm_params.model, while clients use the model_name alias you defined. Direct provider APIs use their own native names. The Swfte gateway uses provider:model, and the model directory shows the id for each model.

    Then map the request fields. OpenRouter accepts a provider object (order, allow_fallbacks, only, ignore, sort, zdr, data_collection, require_parameters) and a models array for fallbacks. These are OpenRouter features. They do not exist on a direct provider API, and on a LiteLLM or Swfte route you re-express them as gateway configuration. Make a row for each one you use.

    Optional OpenRouter headers (HTTP-Referer and X-OpenRouter-Title) only attribute your app on its leaderboards. Remove them when you leave unless the destination asks for something similar.

    Identifier forms by destination
    DestinationModel identifier formWhere gateway features live
    OpenRouterprovider/model or an alias such as ~provider/name-latestprovider object and models array in the request
    LiteLLM proxyClient uses your model_name; config maps it to provider/modelfallbacks, num_retries and keys in config.yaml
    Direct provider (OpenAI-compatible endpoint)The provider's own model nameNothing: you build routing and fallback yourself
    Swfte Connectprovider:modelRouting and failover settings in the platform

    Checked against: OpenRouter quickstart, OpenRouter provider routing, OpenRouter model fallbacks, LiteLLM providers documentation, Swfte alternatives pages (gateway endpoint and model id form)

  3. Step 3Save a replay set of real requests

    You end up with: A file of 100 to 300 real requests, with personal data removed, that you can send to any route.

    The only honest answer to "does it behave the same?" is to send the same requests down both routes. Export requests from your logs, sample across applications and features, and make sure the sample includes your hard cases: the longest prompts, requests with tools, requests that expect JSON, requests in each language you serve, and a handful that currently fail.

    Remove or replace personal data before you save the file, and treat the file as sensitive anyway. Keep one JSON object per line with the messages, the options you send (tools, response format, temperature), and a label for the feature it comes from. Do not store the old route's answers as the truth: they are one sample of a non-deterministic system. You compare behaviour, not exact text.

    If you do not log request bodies today, you will have to write representative requests by hand. That is slower and less honest, and it is a good reason to add logging with redaction as part of this project.

    replay.jsonl (one request per line; write real ones) · json
    {"feature": "support-reply", "model": "<model alias>", "messages": [{"role": "user", "content": "Reply in two sentences: a customer asks for a refund after 45 days."}], "stream": false}
    {"feature": "extract-json", "model": "<model alias>", "messages": [{"role": "user", "content": "Return JSON with keys name and date for: Ana booked on 3 May."}], "stream": false}
  4. Step 4Compare old and new routes on the replay set

    You end up with: A table of differences in success, finish reason, tool calls, JSON validity, token usage and latency between the two routes.

    Write a small script that sends every replay request to the old route and the new route and records what you care about: whether the call succeeded, the HTTP error if not, the finish_reason, whether tool calls came back, whether JSON output parses, the token counts the route reports, and the time taken. The script below does that with the OpenAI SDK pointed at two base URLs. Run it with identical settings and at temperature 0 where the destination allows it.

    Read the differences, not the average. The failures that hurt are specific: a tool call that comes back malformed, a JSON answer wrapped in prose, a request rejected because the destination does not support a parameter, a missing usage field that breaks your cost tracking. Fix each one or decide it is acceptable, and write the decision down.

    Expect some differences in wording. They are not failures unless a downstream rule depends on the exact text. If it does, that rule is fragile and now is a good time to fix it.

    parity.py · python
    import json
    import os
    import sys
    import time
    
    from openai import OpenAI
    
    OLD = OpenAI(base_url=os.environ["OLD_BASE_URL"], api_key=os.environ["OLD_API_KEY"])
    NEW = OpenAI(base_url=os.environ["NEW_BASE_URL"], api_key=os.environ["NEW_API_KEY"])
    MODEL_MAP = json.loads(os.environ.get("MODEL_MAP", "{}"))  # old model id -> new model id
    
    
    def call(client, model, req):
        started = time.perf_counter()
        try:
            r = client.chat.completions.create(
                model=model,
                messages=req["messages"],
                temperature=0,
                max_tokens=req.get("max_tokens", 400),
                **({"tools": req["tools"]} if "tools" in req else {}),
            )
        except Exception as exc:  # record the failure, keep going
            return {"ok": False, "error": f"{type(exc).__name__}: {exc}"[:200]}
        choice = r.choices[0]
        text = choice.message.content or ""
        return {
            "ok": True,
            "finish": choice.finish_reason,
            "tool_calls": len(choice.message.tool_calls or []),
            "json_ok": _is_json(text),
            "usage": r.usage.model_dump() if r.usage else None,
            "seconds": round(time.perf_counter() - started, 2),
        }
    
    
    def _is_json(text):
        try:
            json.loads(text)
            return True
        except ValueError:
            return False
    
    
    for line in open(sys.argv[1], encoding="utf-8"):
        req = json.loads(line)
        old = call(OLD, req["model"], req)
        new = call(NEW, MODEL_MAP.get(req["model"], req["model"]), req)
        diffs = [k for k in ("ok", "finish", "tool_calls", "json_ok") if old.get(k) != new.get(k)]
        print(json.dumps({"feature": req["feature"], "differs_on": diffs, "old": old, "new": new}))
    Run it · bash
    export OLD_BASE_URL="https://openrouter.ai/api/v1"
    export OLD_API_KEY="$OPENROUTER_API_KEY"
    export NEW_BASE_URL="<destination base url>"
    export NEW_API_KEY="<destination key>"
    export MODEL_MAP='{"<old id>": "<new id>"}'
    python parity.py replay.jsonl > parity.jsonl

    Checked against: OpenRouter quickstart

  5. Step 5Change the base URL, key and model names

    You end up with: The client code points at the destination behind a configuration switch, with the old route still available.

    With the parity results in hand, change the client. For any OpenAI-compatible destination the change is three values: the base URL, the API key and the model identifier. Put them in configuration, not in code, so the same build can talk to either route. OpenRouter's base URL is https://openrouter.ai/api/v1. A LiteLLM proxy is whatever address you deployed it at, and uses your key and your model_name aliases. Swfte's gateway endpoint is https://api.swfte.com/agents/v2/gateway, with provider:model ids.

    If your destination is a direct provider, read its compatibility notes first. Anthropic documents an OpenAI SDK compatibility layer at https://api.anthropic.com/v1/, and says it is mainly intended for testing and comparing models, not as a long-term production solution for most use cases. It lists differences that matter: the strict parameter for function calling is ignored, so tool-call JSON is not guaranteed to follow your schema; response_format is ignored; prompt caching is not supported through it; system and developer messages are hoisted and joined into one; n must be 1; temperature above 1 is capped at 1; and many unsupported fields are silently ignored rather than rejected. For production use of Claude directly, use Anthropic's own SDK and API.

    Silent ignoring is the dangerous part. A request that used to enforce a JSON schema through the gateway may now return free text with no error. Your parity test should have caught it; if it did not, add a case for it now.

    Configuration-driven client (Python) · python
    import os
    
    from openai import OpenAI
    
    client = OpenAI(
        base_url=os.environ["LLM_BASE_URL"],  # was https://openrouter.ai/api/v1
        api_key=os.environ["LLM_API_KEY"],
    )
    MODEL = os.environ["LLM_MODEL"]  # use the destination's identifier form
    Test one request against the Swfte gateway endpoint (from the Swfte alternatives pages) · bash
    curl https://api.swfte.com/agents/v2/gateway/chat/completions \
      -H "Authorization: Bearer $SWFTE_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "openai:gpt-4o",
        "messages": [{"role": "user", "content": "Hello"}]
      }'

    Checked against: OpenRouter quickstart, Claude API: OpenAI SDK compatibility, Swfte alternatives pages (gateway endpoint and model id form)

  6. Step 6Re-express the routing and privacy settings you relied on

    You end up with: Each OpenRouter or LiteLLM routing rule has an equivalent on the destination, or a recorded decision to drop it.

    List what the old route did for you without being asked. On OpenRouter, a models array falls back to the next model when the first fails, with the documented triggers of context-length errors, moderation flags on filtered models, rate limiting and downtime, and the request is priced at the model actually used, reported in the model field of the response. The provider object can restrict providers, set an order, prefer low price, throughput or latency, or send only to zero-data-retention endpoints. In-region routing for the EU or US is described as available on its Business and Enterprise plans.

    On a LiteLLM proxy the equivalents are fallbacks, num_retries, allowed_fails and cooldown_time in the config file, and per-key model allow-lists. On other gateways the setting names differ. For every rule, write the destination's equivalent or write "dropped, because ...". Privacy rules need extra care: if you relied on zdr or data_collection: deny, confirm the destination can give you the same guarantee in writing, not only a similar setting name.

    If you move to a single provider, cross-provider fallback and price-based routing are gone. If you move to a self-hosted gateway, you inherit the operation of those features; see how to set up an LLM gateway.

    Examples of settings to re-express
    If you used this on OpenRouterIt did thisWhere to look on the destination
    models: [a, b]Fall back to the next model on listed errorsGateway fallback list, or your own retry code
    provider.order, only, ignoreChoose and exclude upstream providersGateway routing rules, or the provider you chose
    provider.sort by price, throughput, latencyPrefer a cheaper or faster upstreamGateway routing strategy; may not exist
    provider.zdr, data_collection: denyRestrict to zero-data-retention or non-collecting endpointsContract terms and settings; get it in writing
    eu.openrouter.ai endpointIn-region routing (Business and Enterprise plans)Destination's region options and contract

    Checked against: OpenRouter provider routing, OpenRouter model fallbacks, LiteLLM fallbacks and reliability documentation

  7. Step 7Test streaming, tool calls and structured output on their own

    You end up with: Each of the three features has a passing test on the new route, written down as a regression test.

    Streaming is where hand-written parsers break. OpenRouter documents that it periodically sends Server-Sent Events comments such as : OPENROUTER PROCESSING to prevent timeouts, tells clients to skip lines that start with a colon before parsing JSON, ends every Chat Completions stream with an extra chunk that carries usage, and, once headers are sent, reports mid-stream errors as events with an error object and finish_reason: "error". If your parser handles those, check that the destination does not need different handling. If you used the official SDK, you are likely fine, but test.

    For tools, test a call that must use a tool, a call that must not, and a call with several tools. Compare the shape of tool_calls, argument JSON validity and how your code handles a refusal. For structured output, check whether the destination enforces your schema or only asks for it, using a case that tempts the model to add commentary.

    Turn each of these into a permanent test that runs against whichever route is live. You will need them again at the next migration.

    Checked against: OpenRouter streaming, Claude API: OpenAI SDK compatibility

  8. Step 8Match cost tracking and logging

    You end up with: Spend, usage and audit data from the new route lands where your dashboards and alerts expect it.

    Compare the parity run's usage fields. Some routes fill details such as cached tokens and reasoning tokens, and some leave them empty: Anthropic's compatibility layer documents usage.prompt_tokens_details and usage.completion_tokens_details as always empty. If your cost model reads those fields, it will under-report on the new route until you change it.

    Then compare price. Hosted gateways typically pass provider list prices through and add a fee. Our OpenRouter alternatives page records a percentage fee on credit purchases that depends on the plan, so check the current figures on OpenRouter's pricing page against your volume before you decide the migration saves money. A direct route removes that fee but also removes the single invoice and the cross-provider fallback. Work out the monthly difference from your own inventory, with the arithmetic shown, rather than from a comparison page. Our cost reduction guide shows how.

    Finally, logging and audit. Move the alerts, budgets and dashboards that depended on the old gateway's activity data, and decide how long you keep the old data. Switch off the old billing last, and only after the final invoice has matched your records.

    Checked against: Claude API: OpenAI SDK compatibility

  9. Step 9Canary the traffic and keep a rollback

    You end up with: A growing share of live traffic on the new route, with a tested one-step way back.

    Send a small share of real traffic, such as 1 per cent, to the new route, and raise it in steps while you watch. Make the choice sticky per user or session so one conversation does not switch route halfway. Keep the old route configured and warm until the end. A configuration flag that flips back in one step is your rollback; test it before the canary starts, not during an incident.

    Decide the numbers that stop the rollout before you start: for example error rate above the old route's by a set margin, p95 latency above a limit, a drop in your task pass rate, or spend per request above your estimate. Watch them at each step for long enough to cover a full cycle of your traffic, such as a weekday and a weekend if usage differs.

    When you reach 100 per cent, leave the old route in place for a defined period, then remove it, delete unused keys, and close the old account or plan. Update the documents that name the old gateway. Record the migration, the evidence and the date, so the next person can see why the decision was made.

    Sticky percentage canary (Python) · python
    import hashlib
    import os
    
    
    def use_new_route(user_id: str) -> bool:
        """Same user always gets the same route for a given percentage."""
        percent = int(os.environ.get("NEW_ROUTE_PERCENT", "0"))  # 0 is the rollback
        bucket = int(hashlib.sha256(user_id.encode("utf-8")).hexdigest(), 16) % 100
        return bucket < percent

Reasons that do justify it

Data location and control: you need traffic to stay in a region or on your own network. Cost at volume: your spend is large enough that a fee matters, and you use one or two models. Consolidation: three teams run three gateways. A feature you need: budgets per team, guardrails, audit logs. Vendor risk: you want an exit plan. Each of these has a measurable test in the steps above. See our alternatives pages for OpenRouter and LiteLLM for side-by-side facts, each dated.

Troubleshooting

What you seeLikely causeFix
404 or "model not found" after switching the base URLThe model identifier is still in the old gateway's form, such as provider/model.Map every id to the destination's form (step 2). On Swfte use provider:model; on LiteLLM use your model_name alias.
JSON responses arrive as prose, or schemas stop being enforcedThe destination ignores response_format or the strict flag. Anthropic's compatibility layer documents both as ignored.Use the destination's native structured-output feature, or validate and retry in your code. Add a regression test.
Streaming parser throws on the first lines of the streamThe stream contains SSE comment lines such as : OPENROUTER PROCESSING, or differs in how it ends and reports usage.Skip lines that start with a colon before parsing, or use the official SDK. Test the destination's stream on its own.
Cost dashboards drop to zero or under-report after the switchThe usage fields your dashboard reads are empty or named differently on the new route.Compare usage from the parity run and update the cost code. Compute cost from prompt and completion tokens as a fallback.
Tool calls fail or loop on the new routeArgument JSON is not guaranteed to match your schema, or the model behaves differently with your tool descriptions.Validate arguments server-side, return errors to the model as tool results, and compare tool-call tests on both routes.
Requests succeed, but quality complaints rise after the canary growsDifferent default parameters, a different model version behind a name, or a prompt tuned to the old model.Pin model versions, compare sampling parameters, and rerun your task set. Roll back with the flag while you investigate.

Verify it worked

Next steps

Related guides

Frequently asked questions

How do I migrate from OpenRouter?

Inventory the models and features you use, map model ids and routing settings to the destination, replay real requests against both routes, change the base URL, key and model names behind a flag, move traffic in steps with a rollback ready, and match cost tracking before you close the account.

Is migrating just a base URL change?

The code change is small, and the behaviour change is not. Model ids, tool calls, JSON output, streaming, usage fields, routing rules and privacy settings can all differ. A replay test shows which of those affect your application before users do.

What is the OpenRouter base URL and model id format?

The base URL is https://openrouter.ai/api/v1. Model ids take the form provider/model, or an alias such as ~openai/gpt-sol-latest from its quickstart, and the catalogue lists the exact identifiers. Optional headers HTTP-Referer and X-OpenRouter-Title only attribute your app on its leaderboards.

Can I use the OpenAI SDK with Claude directly?

Yes, through Anthropic's compatibility layer at https://api.anthropic.com/v1/. Anthropic says it is mainly for testing and comparison, and lists differences: strict function-calling is ignored, response_format is ignored, prompt caching is not supported, and system messages are joined. For production, use the native API.

How do I replace OpenRouter provider fallbacks?

OpenRouter falls back through a models array or its provider settings. On LiteLLM use the fallbacks list with retries and cooldown in config.yaml. On a single provider you write retry and fallback logic yourself. Test it by stopping the primary route on purpose.

Should I move from LiteLLM to a managed gateway?

Move if running the proxy costs more in attention than it saves, or you need features it lacks. Stay if control, data location and no per-request fee matter more. List what you operate today, including upgrades and key rotation, and compare that with the managed price.

How Swfte can help

You can complete this migration to any destination without Swfte. If Swfte Connect is on your shortlist, it is a managed OpenAI-compatible gateway with routing, failover, cost tracking and budgets, and our alternatives pages set out the comparison with dates.

Connect is a managed service. SAML SSO and SCIM are in development, and dedicated or private deployment is scoped as an engagement, not self-serve. If you need either today, a self-hosted gateway or a direct route may fit better. Supported provider count: <provider count - founder to fill>.

Missing a step or found a command that no longer works? Tell us, or request a how-to.

Sources and last verified

Commands, versions and facts in this guide were checked against the sources below on . Tools change quickly: if something differs from what you see, trust the official documentation and let us know.

  1. OpenRouter quickstart: Base URL https://openrouter.ai/api/v1, OpenAI SDK setup, HTTP-Referer and X-OpenRouter-Title headers, model id format, curl example
  2. OpenRouter provider routing: provider object fields (order, allow_fallbacks, require_parameters, data_collection, zdr, only, ignore, sort) and EU and US in-region routing on Business and Enterprise plans
  3. OpenRouter model fallbacks: models array fallback, triggers, and the model field reporting which model was used and priced
  4. OpenRouter streaming: SSE comment payloads, usage chunk at the end of the stream, mid-stream error format
  5. LiteLLM providers documentation: provider/model naming and the openai/ prefix with api_base for OpenAI-compatible endpoints
  6. LiteLLM fallbacks and reliability documentation: fallbacks, num_retries, allowed_fails and cooldown_time
  7. Claude API: OpenAI SDK compatibility: Base URL, intended use, and the list of ignored or unsupported fields including strict, response_format, prompt caching, system message hoisting, n, temperature cap and empty usage details
  8. Swfte alternatives pages (gateway endpoint and model id form): Swfte gateway endpoint, provider:model id form and the migration outline, as published on swfte.com

Topics

  • migration
  • OpenRouter
  • LiteLLM
  • LLM gateway
  • parity testing

Machine-readable copies: this guide as markdown, index of all guides (JSON). Canonical address: https://www.swfte.com/how-to-migrate-from-openrouter-or-litellm.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.