Short answer
Changing a base URL takes two lines, and the risk is in everything else: model names, tool calls, structured output, streaming, usage fields, routing rules and cost. Inventory what you use, save a replay set of real requests, compare old and new routes on it, switch the client behind a percentage flag, watch cost and errors, and keep the old route until you are sure. Do not migrate if your reason is only habit.
The steps at a glance
- Inventory what you actually use
- Map model names and request fields
- Save a replay set of real requests
- Compare old and new routes on the replay set
- Change the base URL, key and model names
- Re-express the routing and privacy settings you relied on
- Test streaming, tool calls and structured output on their own
- Match cost tracking and logging
- Canary the traffic and keep a rollback
Before you start
Who this is for
- Engineers whose application calls OpenRouter or a LiteLLM proxy and who have decided to move to direct provider APIs, another gateway, or Swfte Connect.
- Platform teams consolidating several gateways, or moving from a hosted gateway to one they run.
- Anyone asked "can we just change the URL?" who wants the evidence behind the answer.
Probably not for you if
- Teams building a first gateway. Start with how to set up an LLM gateway.
- Teams whose current setup works and whose only reason to move is novelty. Read the next section before you spend a week on this.
Prerequisites
- Access to the code and configuration that call the current gateway, and the ability to deploy a change behind a flag.
- Keys or accounts for the destination, and permission to send test traffic that costs money.
- A way to read your current usage: request logs, the gateway's activity export, or the provider dashboard.
- Python 3 and the
openaipackage for the parity script (pip install openai). - Agreement on who approves the cut-over and who can roll it back.
- Time
- About 1 to 2 days of work for one application, plus a canary period of a few days
- Cost
- The cost of replaying a few hundred requests on both routes, plus any overlap in provider spend during the canary.
- Skill
- Comfortable changing production API client code and reading usage logs
Estimates are ours, not measurements, and move with your hardware, data and network.
When not to migrate
Most migration pages on this topic are written by the gateway that wants your traffic, and they describe the move as two lines of code. Two lines are the cheap part. Before you start, check whether your reason holds up.
- You use several providers through one key and value that. Moving to one provider's API gives up the single invoice and the fallback across providers. OpenRouter's breadth is its point.
- The saving is small against the engineering time. Our OpenRouter against direct Anthropic comparison frames the fee as the price of flexibility. Do your own arithmetic with your volume.
- You would be trading a service you do not operate for software you must. LiteLLM gives you control, and you carry upgrades, security and uptime. See how to set up an LLM gateway for what that involves, including the March 2026 supply-chain incident.
- You rely on a feature the destination lacks. For example, request-level provider routing, zero-data-retention routing, or a schema guarantee.
- Nothing is wrong today. A working route is an asset. Migrate for a reason you can name: cost, control, data location, a feature, or consolidation.
Step 1Inventory what you actually use
You end up with: A list of every model, feature and gateway-specific field your applications rely on, with a rough volume for each.
Start from facts, not from memory. Search the code for the gateway's base URL and for any code that adds gateway-specific fields to requests. Then pull a month of usage from the gateway or provider and group it by model, by application and by feature: streaming, tool calls, JSON output, image input, long prompts.
Write the result in a table with one row per model and application. The columns that matter most are request volume, whether the call streams, whether it uses tools, whether it relies on JSON output, and whether it uses any routing preference. Volume tells you where a saving or a mistake is large. The feature columns tell you where the destination might behave differently.
Also list the non-code dependencies: dashboards that read the gateway's activity logs, budget alerts, and any contract or data processing terms. A migration that forgets the finance dashboard fails in month two.
Find calls to the current gateway · bash grep -rn "openrouter.ai" --include="*.py" --include="*.ts" --include="*.js" --include="*.yaml" --include="*.json" . grep -rn "OPENROUTER_API_KEY\|LITELLM" --include="*.py" --include="*.ts" --include="*.js" --include="*.yaml" --include="*.env.example" .Inventory template Application Model Requests per month Streams? Tools? JSON output? Routing preference? support-bot record the model id as sent today from usage export yes or no yes or no yes or no provider order, fallbacks, zdr Step 2Map model names and request fields
You end up with: For every model in the inventory, the exact identifier the destination expects, and a note for each field that does not carry over.
Model names are the first thing to break. OpenRouter model ids are of the form
provider/model, for example the~openai/gpt-sol-lateststyle alias its quickstart uses, or a specific identifier from its model catalogue. LiteLLM also names modelsprovider/model, but there that string goes inlitellm_params.model, while clients use themodel_namealias you defined. Direct provider APIs use their own native names. The Swfte gateway usesprovider:model, and the model directory shows the id for each model.Then map the request fields. OpenRouter accepts a
providerobject (order,allow_fallbacks,only,ignore,sort,zdr,data_collection,require_parameters) and amodelsarray for fallbacks. These are OpenRouter features. They do not exist on a direct provider API, and on a LiteLLM or Swfte route you re-express them as gateway configuration. Make a row for each one you use.Optional OpenRouter headers (
HTTP-RefererandX-OpenRouter-Title) only attribute your app on its leaderboards. Remove them when you leave unless the destination asks for something similar.Identifier forms by destination Destination Model identifier form Where gateway features live OpenRouter provider/modelor an alias such as~provider/name-latestproviderobject andmodelsarray in the requestLiteLLM proxy Client uses your model_name; config maps it toprovider/modelfallbacks,num_retriesand keys inconfig.yamlDirect provider (OpenAI-compatible endpoint) The provider's own model name Nothing: you build routing and fallback yourself Swfte Connect provider:modelRouting and failover settings in the platform Checked against: OpenRouter quickstart, OpenRouter provider routing, OpenRouter model fallbacks, LiteLLM providers documentation, Swfte alternatives pages (gateway endpoint and model id form)
Step 3Save a replay set of real requests
You end up with: A file of 100 to 300 real requests, with personal data removed, that you can send to any route.
The only honest answer to "does it behave the same?" is to send the same requests down both routes. Export requests from your logs, sample across applications and features, and make sure the sample includes your hard cases: the longest prompts, requests with tools, requests that expect JSON, requests in each language you serve, and a handful that currently fail.
Remove or replace personal data before you save the file, and treat the file as sensitive anyway. Keep one JSON object per line with the messages, the options you send (tools, response format, temperature), and a label for the feature it comes from. Do not store the old route's answers as the truth: they are one sample of a non-deterministic system. You compare behaviour, not exact text.
If you do not log request bodies today, you will have to write representative requests by hand. That is slower and less honest, and it is a good reason to add logging with redaction as part of this project.
replay.jsonl (one request per line; write real ones) · json {"feature": "support-reply", "model": "<model alias>", "messages": [{"role": "user", "content": "Reply in two sentences: a customer asks for a refund after 45 days."}], "stream": false} {"feature": "extract-json", "model": "<model alias>", "messages": [{"role": "user", "content": "Return JSON with keys name and date for: Ana booked on 3 May."}], "stream": false}Step 4Compare old and new routes on the replay set
You end up with: A table of differences in success, finish reason, tool calls, JSON validity, token usage and latency between the two routes.
Write a small script that sends every replay request to the old route and the new route and records what you care about: whether the call succeeded, the HTTP error if not, the
finish_reason, whether tool calls came back, whether JSON output parses, the token counts the route reports, and the time taken. The script below does that with the OpenAI SDK pointed at two base URLs. Run it with identical settings and at temperature 0 where the destination allows it.Read the differences, not the average. The failures that hurt are specific: a tool call that comes back malformed, a JSON answer wrapped in prose, a request rejected because the destination does not support a parameter, a missing usage field that breaks your cost tracking. Fix each one or decide it is acceptable, and write the decision down.
Expect some differences in wording. They are not failures unless a downstream rule depends on the exact text. If it does, that rule is fragile and now is a good time to fix it.
parity.py · python import json import os import sys import time from openai import OpenAI OLD = OpenAI(base_url=os.environ["OLD_BASE_URL"], api_key=os.environ["OLD_API_KEY"]) NEW = OpenAI(base_url=os.environ["NEW_BASE_URL"], api_key=os.environ["NEW_API_KEY"]) MODEL_MAP = json.loads(os.environ.get("MODEL_MAP", "{}")) # old model id -> new model id def call(client, model, req): started = time.perf_counter() try: r = client.chat.completions.create( model=model, messages=req["messages"], temperature=0, max_tokens=req.get("max_tokens", 400), **({"tools": req["tools"]} if "tools" in req else {}), ) except Exception as exc: # record the failure, keep going return {"ok": False, "error": f"{type(exc).__name__}: {exc}"[:200]} choice = r.choices[0] text = choice.message.content or "" return { "ok": True, "finish": choice.finish_reason, "tool_calls": len(choice.message.tool_calls or []), "json_ok": _is_json(text), "usage": r.usage.model_dump() if r.usage else None, "seconds": round(time.perf_counter() - started, 2), } def _is_json(text): try: json.loads(text) return True except ValueError: return False for line in open(sys.argv[1], encoding="utf-8"): req = json.loads(line) old = call(OLD, req["model"], req) new = call(NEW, MODEL_MAP.get(req["model"], req["model"]), req) diffs = [k for k in ("ok", "finish", "tool_calls", "json_ok") if old.get(k) != new.get(k)] print(json.dumps({"feature": req["feature"], "differs_on": diffs, "old": old, "new": new}))Run it · bash export OLD_BASE_URL="https://openrouter.ai/api/v1" export OLD_API_KEY="$OPENROUTER_API_KEY" export NEW_BASE_URL="<destination base url>" export NEW_API_KEY="<destination key>" export MODEL_MAP='{"<old id>": "<new id>"}' python parity.py replay.jsonl > parity.jsonlChecked against: OpenRouter quickstart
Step 5Change the base URL, key and model names
You end up with: The client code points at the destination behind a configuration switch, with the old route still available.
With the parity results in hand, change the client. For any OpenAI-compatible destination the change is three values: the base URL, the API key and the model identifier. Put them in configuration, not in code, so the same build can talk to either route. OpenRouter's base URL is
https://openrouter.ai/api/v1. A LiteLLM proxy is whatever address you deployed it at, and uses your key and yourmodel_namealiases. Swfte's gateway endpoint ishttps://api.swfte.com/agents/v2/gateway, withprovider:modelids.If your destination is a direct provider, read its compatibility notes first. Anthropic documents an OpenAI SDK compatibility layer at
https://api.anthropic.com/v1/, and says it is mainly intended for testing and comparing models, not as a long-term production solution for most use cases. It lists differences that matter: thestrictparameter for function calling is ignored, so tool-call JSON is not guaranteed to follow your schema;response_formatis ignored; prompt caching is not supported through it; system and developer messages are hoisted and joined into one;nmust be 1;temperatureabove 1 is capped at 1; and many unsupported fields are silently ignored rather than rejected. For production use of Claude directly, use Anthropic's own SDK and API.Silent ignoring is the dangerous part. A request that used to enforce a JSON schema through the gateway may now return free text with no error. Your parity test should have caught it; if it did not, add a case for it now.
Configuration-driven client (Python) · python import os from openai import OpenAI client = OpenAI( base_url=os.environ["LLM_BASE_URL"], # was https://openrouter.ai/api/v1 api_key=os.environ["LLM_API_KEY"], ) MODEL = os.environ["LLM_MODEL"] # use the destination's identifier formTest one request against the Swfte gateway endpoint (from the Swfte alternatives pages) · bash curl https://api.swfte.com/agents/v2/gateway/chat/completions \ -H "Authorization: Bearer $SWFTE_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "openai:gpt-4o", "messages": [{"role": "user", "content": "Hello"}] }'Checked against: OpenRouter quickstart, Claude API: OpenAI SDK compatibility, Swfte alternatives pages (gateway endpoint and model id form)
Step 6Re-express the routing and privacy settings you relied on
You end up with: Each OpenRouter or LiteLLM routing rule has an equivalent on the destination, or a recorded decision to drop it.
List what the old route did for you without being asked. On OpenRouter, a
modelsarray falls back to the next model when the first fails, with the documented triggers of context-length errors, moderation flags on filtered models, rate limiting and downtime, and the request is priced at the model actually used, reported in themodelfield of the response. Theproviderobject can restrict providers, set an order, prefer low price, throughput or latency, or send only to zero-data-retention endpoints. In-region routing for the EU or US is described as available on its Business and Enterprise plans.On a LiteLLM proxy the equivalents are
fallbacks,num_retries,allowed_failsandcooldown_timein the config file, and per-key model allow-lists. On other gateways the setting names differ. For every rule, write the destination's equivalent or write "dropped, because ...". Privacy rules need extra care: if you relied onzdrordata_collection: deny, confirm the destination can give you the same guarantee in writing, not only a similar setting name.If you move to a single provider, cross-provider fallback and price-based routing are gone. If you move to a self-hosted gateway, you inherit the operation of those features; see how to set up an LLM gateway.
Examples of settings to re-express If you used this on OpenRouter It did this Where to look on the destination models: [a, b]Fall back to the next model on listed errors Gateway fallback list, or your own retry code provider.order,only,ignoreChoose and exclude upstream providers Gateway routing rules, or the provider you chose provider.sortby price, throughput, latencyPrefer a cheaper or faster upstream Gateway routing strategy; may not exist provider.zdr,data_collection: denyRestrict to zero-data-retention or non-collecting endpoints Contract terms and settings; get it in writing eu.openrouter.aiendpointIn-region routing (Business and Enterprise plans) Destination's region options and contract Checked against: OpenRouter provider routing, OpenRouter model fallbacks, LiteLLM fallbacks and reliability documentation
Step 7Test streaming, tool calls and structured output on their own
You end up with: Each of the three features has a passing test on the new route, written down as a regression test.
Streaming is where hand-written parsers break. OpenRouter documents that it periodically sends Server-Sent Events comments such as
: OPENROUTER PROCESSINGto prevent timeouts, tells clients to skip lines that start with a colon before parsing JSON, ends every Chat Completions stream with an extra chunk that carriesusage, and, once headers are sent, reports mid-stream errors as events with anerrorobject andfinish_reason: "error". If your parser handles those, check that the destination does not need different handling. If you used the official SDK, you are likely fine, but test.For tools, test a call that must use a tool, a call that must not, and a call with several tools. Compare the shape of
tool_calls, argument JSON validity and how your code handles a refusal. For structured output, check whether the destination enforces your schema or only asks for it, using a case that tempts the model to add commentary.Turn each of these into a permanent test that runs against whichever route is live. You will need them again at the next migration.
Checked against: OpenRouter streaming, Claude API: OpenAI SDK compatibility
Step 8Match cost tracking and logging
You end up with: Spend, usage and audit data from the new route lands where your dashboards and alerts expect it.
Compare the parity run's
usagefields. Some routes fill details such as cached tokens and reasoning tokens, and some leave them empty: Anthropic's compatibility layer documentsusage.prompt_tokens_detailsandusage.completion_tokens_detailsas always empty. If your cost model reads those fields, it will under-report on the new route until you change it.Then compare price. Hosted gateways typically pass provider list prices through and add a fee. Our OpenRouter alternatives page records a percentage fee on credit purchases that depends on the plan, so check the current figures on OpenRouter's pricing page against your volume before you decide the migration saves money. A direct route removes that fee but also removes the single invoice and the cross-provider fallback. Work out the monthly difference from your own inventory, with the arithmetic shown, rather than from a comparison page. Our cost reduction guide shows how.
Finally, logging and audit. Move the alerts, budgets and dashboards that depended on the old gateway's activity data, and decide how long you keep the old data. Switch off the old billing last, and only after the final invoice has matched your records.
Checked against: Claude API: OpenAI SDK compatibility
Step 9Canary the traffic and keep a rollback
You end up with: A growing share of live traffic on the new route, with a tested one-step way back.
Send a small share of real traffic, such as 1 per cent, to the new route, and raise it in steps while you watch. Make the choice sticky per user or session so one conversation does not switch route halfway. Keep the old route configured and warm until the end. A configuration flag that flips back in one step is your rollback; test it before the canary starts, not during an incident.
Decide the numbers that stop the rollout before you start: for example error rate above the old route's by a set margin, p95 latency above a limit, a drop in your task pass rate, or spend per request above your estimate. Watch them at each step for long enough to cover a full cycle of your traffic, such as a weekday and a weekend if usage differs.
When you reach 100 per cent, leave the old route in place for a defined period, then remove it, delete unused keys, and close the old account or plan. Update the documents that name the old gateway. Record the migration, the evidence and the date, so the next person can see why the decision was made.
Sticky percentage canary (Python) · python import hashlib import os def use_new_route(user_id: str) -> bool: """Same user always gets the same route for a given percentage.""" percent = int(os.environ.get("NEW_ROUTE_PERCENT", "0")) # 0 is the rollback bucket = int(hashlib.sha256(user_id.encode("utf-8")).hexdigest(), 16) % 100 return bucket < percent
Reasons that do justify it
Data location and control: you need traffic to stay in a region or on your own network. Cost at volume: your spend is large enough that a fee matters, and you use one or two models. Consolidation: three teams run three gateways. A feature you need: budgets per team, guardrails, audit logs. Vendor risk: you want an exit plan. Each of these has a measurable test in the steps above. See our alternatives pages for OpenRouter and LiteLLM for side-by-side facts, each dated.
Troubleshooting
| What you see | Likely cause | Fix |
|---|---|---|
| 404 or "model not found" after switching the base URL | The model identifier is still in the old gateway's form, such as provider/model. | Map every id to the destination's form (step 2). On Swfte use provider:model; on LiteLLM use your model_name alias. |
| JSON responses arrive as prose, or schemas stop being enforced | The destination ignores response_format or the strict flag. Anthropic's compatibility layer documents both as ignored. | Use the destination's native structured-output feature, or validate and retry in your code. Add a regression test. |
| Streaming parser throws on the first lines of the stream | The stream contains SSE comment lines such as : OPENROUTER PROCESSING, or differs in how it ends and reports usage. | Skip lines that start with a colon before parsing, or use the official SDK. Test the destination's stream on its own. |
| Cost dashboards drop to zero or under-report after the switch | The usage fields your dashboard reads are empty or named differently on the new route. | Compare usage from the parity run and update the cost code. Compute cost from prompt and completion tokens as a fallback. |
| Tool calls fail or loop on the new route | Argument JSON is not guaranteed to match your schema, or the model behaves differently with your tool descriptions. | Validate arguments server-side, return errors to the model as tool results, and compare tool-call tests on both routes. |
| Requests succeed, but quality complaints rise after the canary grows | Different default parameters, a different model version behind a name, or a prompt tuned to the old model. | Pin model versions, compare sampling parameters, and rerun your task set. Roll back with the flag while you investigate. |
Verify it worked
Next steps
- How to set up an LLM gateway: run the gateway yourself if that is where you are moving
- How to reduce LLM costs: do the cost arithmetic from your own usage
- OpenRouter alternatives: dated facts on OpenRouter and what Swfte Connect offers
- LiteLLM alternatives: dated facts on LiteLLM and how it compares
- OpenRouter or Anthropic direct: the cost and flexibility trade-off for Claude workloads
Related guides
- How to Set Up an LLM Gateway with LiteLLM (2026): Run the open-source LiteLLM proxy in Docker with a config file, add a model, issue virtual keys with budgets, set fallbacks, wire health checks, and harden it for production.
- How to Reduce LLM Costs: Step-by-Step Guide (2026): Measure token spend per feature first, then apply the levers in order of payoff: output caps, prompt caching, batch APIs, model routing, response caching, and a self-hosting break-even check.
- How to Choose an LLM for Your Company: A Scorecard: A selection process, not a leaderboard: requirements, a hosted and open-weight shortlist, a test on your own tasks, a weighted scorecard, licence and data-terms checks, and an exit plan.
- How to Monitor AI Agents in Production (2026 Guide): Give every agent run an id, record each model and tool step as a span, redact before you store, alert on loops, tool failures and cost per run, and read a weekly sample by hand.
Frequently asked questions
How do I migrate from OpenRouter?
Inventory the models and features you use, map model ids and routing settings to the destination, replay real requests against both routes, change the base URL, key and model names behind a flag, move traffic in steps with a rollback ready, and match cost tracking before you close the account.
Is migrating just a base URL change?
The code change is small, and the behaviour change is not. Model ids, tool calls, JSON output, streaming, usage fields, routing rules and privacy settings can all differ. A replay test shows which of those affect your application before users do.
What is the OpenRouter base URL and model id format?
The base URL is https://openrouter.ai/api/v1. Model ids take the form provider/model, or an alias such as ~openai/gpt-sol-latest from its quickstart, and the catalogue lists the exact identifiers. Optional headers HTTP-Referer and X-OpenRouter-Title only attribute your app on its leaderboards.
Can I use the OpenAI SDK with Claude directly?
Yes, through Anthropic's compatibility layer at https://api.anthropic.com/v1/. Anthropic says it is mainly for testing and comparison, and lists differences: strict function-calling is ignored, response_format is ignored, prompt caching is not supported, and system messages are joined. For production, use the native API.
How do I replace OpenRouter provider fallbacks?
OpenRouter falls back through a models array or its provider settings. On LiteLLM use the fallbacks list with retries and cooldown in config.yaml. On a single provider you write retry and fallback logic yourself. Test it by stopping the primary route on purpose.
Should I move from LiteLLM to a managed gateway?
Move if running the proxy costs more in attention than it saves, or you need features it lacks. Stay if control, data location and no per-request fee matter more. List what you operate today, including upgrades and key rotation, and compare that with the managed price.
How Swfte can help
You can complete this migration to any destination without Swfte. If Swfte Connect is on your shortlist, it is a managed OpenAI-compatible gateway with routing, failover, cost tracking and budgets, and our alternatives pages set out the comparison with dates.
- Swfte Connect: the managed gateway and what it covers
- Moving from OpenRouter to Swfte: the side-by-side facts and a short migration outline
- Moving from LiteLLM to Swfte: what changes if you stop running the proxy yourself
Connect is a managed service. SAML SSO and SCIM are in development, and dedicated or private deployment is scoped as an engagement, not self-serve. If you need either today, a self-hosted gateway or a direct route may fit better. Supported provider count: <provider count - founder to fill>.
Missing a step or found a command that no longer works? Tell us, or request a how-to.
Sources and last verified
Commands, versions and facts in this guide were checked against the sources below on . Tools change quickly: if something differs from what you see, trust the official documentation and let us know.
- OpenRouter quickstart: Base URL https://openrouter.ai/api/v1, OpenAI SDK setup, HTTP-Referer and X-OpenRouter-Title headers, model id format, curl example
- OpenRouter provider routing: provider object fields (order, allow_fallbacks, require_parameters, data_collection, zdr, only, ignore, sort) and EU and US in-region routing on Business and Enterprise plans
- OpenRouter model fallbacks: models array fallback, triggers, and the model field reporting which model was used and priced
- OpenRouter streaming: SSE comment payloads, usage chunk at the end of the stream, mid-stream error format
- LiteLLM providers documentation: provider/model naming and the openai/ prefix with api_base for OpenAI-compatible endpoints
- LiteLLM fallbacks and reliability documentation: fallbacks, num_retries, allowed_fails and cooldown_time
- Claude API: OpenAI SDK compatibility: Base URL, intended use, and the list of ignored or unsupported fields including strict, response_format, prompt caching, system message hoisting, n, temperature cap and empty usage details
- Swfte alternatives pages (gateway endpoint and model id form): Swfte gateway endpoint, provider:model id form and the migration outline, as published on swfte.com
Topics
- migration
- OpenRouter
- LiteLLM
- LLM gateway
- parity testing
Machine-readable copies: this guide as markdown, index of all guides (JSON). Canonical address: https://www.swfte.com/how-to-migrate-from-openrouter-or-litellm.