Operate · Intermediate

How to reduce LLM costs

  • Time: About 3 hours to add measurement and the first three levers; routing and batch work take longer
  • Cost: No new software to buy. Expect to spend a little on test calls while you compare models.
  • Level: Intermediate
On this page
  1. Short answer
  2. Before you start
  3. 1. Log token usage on every model call
  4. 2. Rank features by cost and find the top three
  5. 3. Cap output length and tighten the prompt
  6. 4. Turn on prompt caching for the static start of every prompt
  7. 5. Move work that can wait to a batch API
  8. 6. Route easy requests to a cheaper model, behind an eval gate
  9. 7. Cache whole answers and trim what you retrieve
  10. 8. Set budgets and alerts, then price self-hosting honestly
  11. Worked example: what each lever is worth (hypothetical numbers)
  12. Where cost cutting goes wrong
  13. Troubleshooting
  14. Verify it worked
  15. Next steps
  16. FAQ
  17. How Swfte can help
  18. Sources and last verified

Short answer

Measure tokens per feature, model and user before changing anything. Then work through the levers in order: cap output length, turn on prompt caching for the static start of every prompt, move work that can wait to a batch API at half price, route easy requests to a cheaper model behind an eval gate, and only then price self-hosting. Provider prices and cache rules change, so read each lever against the current docs.

The steps at a glance

  1. Log token usage on every model call
  2. Rank features by cost and find the top three
  3. Cap output length and tighten the prompt
  4. Turn on prompt caching for the static start of every prompt
  5. Move work that can wait to a batch API
  6. Route easy requests to a cheaper model, behind an eval gate
  7. Cache whole answers and trim what you retrieve
  8. Set budgets and alerts, then price self-hosting honestly

Before you start

Who this is for

  • Developers and platform engineers whose monthly model bill has become a line item someone asks about.
  • Technical leads who need a ranked list of savings with the arithmetic shown, not a list of tips.
  • Teams running agents or retrieval pipelines, where a single run can make many model calls.

Probably not for you if

Prerequisites

  • Access to the code that calls the model API, and permission to change it.
  • Python 3.10 or later for the example code (the current Anthropic and OpenAI Python packages require it).
  • An API key for at least one provider, kept in an environment variable and not in code.
  • A set of 30 to 50 real requests with outputs you have judged acceptable. You need this before you try routing or a cheaper model.
Time
About 3 hours to add measurement and the first three levers; routing and batch work take longer
Cost
No new software to buy. Expect to spend a little on test calls while you compare models.
Hardware
None, unless you decide to self-host in the last step.
Skill
Comfortable reading and editing production code that calls an LLM API

Estimates are ours, not measurements, and move with your hardware, data and network.

  1. Step 1Log token usage on every model call

    You end up with: Every call writes one line saying which feature, model and user it belongs to, and how many tokens it used.

    You cannot cut a bill you cannot attribute. Provider dashboards show spend per key, which is rarely the unit you care about. Wrap your model calls in one small function that writes a log line per call with a feature name, a model name, a user or tenant id and the token counts the provider returns.

    The field names differ by provider. Anthropic returns input_tokens, output_tokens, cache_creation_input_tokens and cache_read_input_tokens in the usage object. OpenAI Responses returns input_tokens, output_tokens, input_tokens_details.cached_tokens and output_tokens_details.reasoning_tokens. Log all of them. Reasoning tokens and cache writes are billed, and they are the usual reason a bill does not match a naive count.

    Do not log prompt text here. Token counts and ids are enough for cost work, and keeping content out of the log avoids a second data-protection problem.

    Install the SDKs · bash
    pip install anthropic openai
    Usage logging wrapper (Anthropic and OpenAI Responses) · python
    import json
    import time
    
    from anthropic import Anthropic
    from openai import OpenAI
    
    anthropic_client = Anthropic()  # reads ANTHROPIC_API_KEY
    openai_client = OpenAI()        # reads OPENAI_API_KEY
    
    
    def log_usage(row):
        print(json.dumps(row))  # swap print for your logger or a file
    
    
    def call_claude(feature, user_id, **kwargs):
        start = time.time()
        resp = anthropic_client.messages.create(**kwargs)
        u = resp.usage
        log_usage({
            "feature": feature,
            "user": user_id,
            "model": kwargs["model"],
            "input_tokens": u.input_tokens,
            "output_tokens": u.output_tokens,
            "cache_write_tokens": getattr(u, "cache_creation_input_tokens", 0) or 0,
            "cache_read_tokens": getattr(u, "cache_read_input_tokens", 0) or 0,
            "stop_reason": resp.stop_reason,
            "ms": int((time.time() - start) * 1000),
        })
        return resp
    
    
    def call_openai(feature, user_id, **kwargs):
        start = time.time()
        resp = openai_client.responses.create(**kwargs)
        u = resp.usage
        log_usage({
            "feature": feature,
            "user": user_id,
            "model": kwargs["model"],
            "input_tokens": u.input_tokens,
            "output_tokens": u.output_tokens,
            "cache_read_tokens": u.input_tokens_details.cached_tokens,
            "reasoning_tokens": u.output_tokens_details.reasoning_tokens,
            "ms": int((time.time() - start) * 1000),
        })
        return resp

    Checked against: Anthropic: Messages API reference, OpenAI: Responses API reference, PyPI: anthropic, PyPI: openai

  2. Step 2Rank features by cost and find the top three

    You end up with: You have a ranked list showing which three features, models or users cause most of the spend.

    After a day or a week of logs, add the tokens up by feature and by model. Multiply by your current per-million-token prices to get money. Most teams find that two or three features cause most of the cost, and that within those features a few long prompts or very long outputs matter more than the average request.

    Look for four patterns. A large static prefix repeated on every call (caching will help). Output far longer than the task needs (a cap will help). Work nobody is waiting for, such as nightly enrichment or evaluation runs (a batch API will help). And an expensive model doing simple work such as classification or extraction (routing will help). Write down which pattern each of your top three matches before you touch any code.

    Total tokens by feature and model from the JSON-lines log · python
    import json
    import sys
    from collections import defaultdict
    
    totals = defaultdict(lambda: {"calls": 0, "in": 0, "out": 0, "cache_read": 0})
    for line in open(sys.argv[1]):
        r = json.loads(line)
        t = totals[(r["feature"], r["model"])]
        t["calls"] += 1
        t["in"] += r["input_tokens"]
        t["out"] += r["output_tokens"]
        t["cache_read"] += r.get("cache_read_tokens", 0)
    
    for (feature, model), t in sorted(totals.items(), key=lambda kv: -(kv[1]["in"] + kv[1]["out"])):
        print(f"{feature:24} {model:28} calls={t['calls']:7} in={t['in']:12} out={t['out']:10} cached={t['cache_read']:12}")
    Run it (save the script as rank_spend.py and your log as usage.jsonl) · bash
    python rank_spend.py usage.jsonl
    Match each expensive pattern to a lever
    Pattern in your logLeverStep
    Large identical prefix, many callsPrompt caching4
    Output tokens much larger than the task needsOutput cap and tighter prompt3
    Work nobody is waiting forBatch API5
    Frontier model on classification or extractionModel routing6
    Same question asked repeatedlyResponse cache7
  3. Step 3Cap output length and tighten the prompt

    You end up with: Each feature has an explicit output limit and no longer asks for more text than the task needs.

    Output tokens cost more than input tokens at every major provider, so unbounded output is the cheapest place to save. Set an explicit limit on every call. Anthropic calls it max_tokens, the maximum number of tokens to generate before stopping. OpenAI Responses calls it max_output_tokens, an upper bound that includes visible output and reasoning tokens.

    Set the cap from your log, not from a guess: take the 99th percentile of output tokens for the feature and add a margin. Then check the stop_reason field on Anthropic responses. A value of max_tokens means the answer was cut off, so the cap is too low for that feature. Ask for the format you need in the prompt itself ("answer in two sentences", "return only the JSON object") rather than relying on the cap to truncate.

    Trim the input as well. Remove instructions the model no longer needs, drop examples that never change the answer, and stop re-sending a whole conversation when a short summary of it would do. Keep an eye on quality while you do this, because a shorter prompt that lowers accuracy is not a saving.

    Checked against: Anthropic: Messages API reference, OpenAI: Responses API reference

  4. Step 4Turn on prompt caching for the static start of every prompt

    You end up with: Repeated prefixes are billed at the cheaper cache-read rate, and your log shows cache read tokens above zero.

    Put everything that does not change first (system prompt, tool definitions, reference documents) and everything that changes last (the user message). Caching works on the unchanged beginning of the prompt, so one timestamp or request id near the top defeats it.

    On Anthropic, add cache_control with type ephemeral to the system block, or at the top level of the request for automatic caching. The default lifetime is 5 minutes, and "ttl": "1h" selects a 1-hour cache. As of 2026-10-06 the docs list a 5-minute cache write at 1.25 times the base input price, a 1-hour write at 2 times, and a cache read at 0.1 times for most models, with lower read multipliers on a few newer ones. Prompts shorter than the model minimum (512, 1,024, 2,048 or 4,096 tokens depending on the model) are not cached.

    On OpenAI, caching is on by default for supported models, so there is nothing to switch on. The docs give a 1,024-token minimum on the newest models, a 30-minute default lifetime there, and a read price of 0.1 times the uncached input price. Earlier models use prompt_cache_retention with in_memory or 24h. A stable prompt_cache_key helps requests with a shared prefix land together. Check the hit rate in the field you logged: cached tokens divided by total input tokens.

    Anthropic: mark the static system prompt as cacheable · python
    LONG_STATIC_PROMPT = open("system_prompt.txt").read()  # must exceed the model minimum
    
    resp = call_claude(
        "support-triage",
        "user-123",
        model="YOUR_MODEL_ID",
        max_tokens=400,
        system=[
            {
                "type": "text",
                "text": LONG_STATIC_PROMPT,
                "cache_control": {"type": "ephemeral"},
            }
        ],
        messages=[{"role": "user", "content": "Customer says the export button does nothing."}],
    )
    Anthropic: 1-hour cache for prompts used less often than every 5 minutes · json
    {"cache_control": {"type": "ephemeral", "ttl": "1h"}}

    Second call within the cache lifetime, from your log

    {"feature": "support-triage", "cache_write_tokens": 0, "cache_read_tokens": <size of the cached prefix>, ...}

    Checked against: Anthropic: Prompt caching, OpenAI: Prompt caching

  5. Step 5Move work that can wait to a batch API

    You end up with: Offline jobs run through a batch endpoint at half the standard price.

    Evaluation runs, nightly enrichment, bulk classification and back-fills do not need an answer in two seconds. Both Anthropic and OpenAI document a 50% discount for batch processing. On Anthropic, all batch usage is charged at 50% of standard prices. On OpenAI, the batch guide states a 50% discount and a completion window that can currently only be set to 24h.

    The limits matter when you plan the job. Anthropic allows up to 100,000 requests or 256 MB per batch, most batches finish within an hour, they expire after 24 hours, and results stay available for 29 days. OpenAI allows up to 50,000 requests and a 200 MB input file per batch. Each request carries a custom_id so you can match results to inputs.

    The example below creates an Anthropic batch with two requests and then polls and reads the results. Streaming is not supported inside a batch. Check the provider docs for whether the batch discount combines with caching on your model, because Anthropic suggests the 1-hour cache for batches that share context.

    Anthropic: create a batch · python
    import anthropic
    from anthropic.types.message_create_params import MessageCreateParamsNonStreaming
    from anthropic.types.messages.batch_create_params import Request
    
    client = anthropic.Anthropic()
    
    message_batch = client.messages.batches.create(
        requests=[
            Request(
                custom_id="doc-0001",
                params=MessageCreateParamsNonStreaming(
                    model="YOUR_MODEL_ID",
                    max_tokens=300,
                    messages=[{"role": "user", "content": "Classify this ticket: ..."}],
                ),
            ),
            Request(
                custom_id="doc-0002",
                params=MessageCreateParamsNonStreaming(
                    model="YOUR_MODEL_ID",
                    max_tokens=300,
                    messages=[{"role": "user", "content": "Classify this ticket: ..."}],
                ),
            ),
        ]
    )
    print(message_batch.id)
    Anthropic: poll until the batch has ended, then stream the results · python
    import time
    
    while True:
        batch = client.messages.batches.retrieve(message_batch.id)
        if batch.processing_status == "ended":
            break
        print("still processing...")
        time.sleep(60)
    
    for result in client.messages.batches.results(message_batch.id):
        outcome = result.result
        match outcome.type:
            case "succeeded":
                print("ok", result.custom_id)
            case "errored":
                print("errored", result.custom_id)
            case "expired":
                print("expired", result.custom_id)
    OpenAI: upload a JSONL file and create a batch · python
    from openai import OpenAI
    
    client = OpenAI()
    
    batch_input_file = client.files.create(
        file=open("batchinput.jsonl", "rb"), purpose="batch"
    )
    
    batch = client.batches.create(
        input_file_id=batch_input_file.id,
        endpoint="/v1/chat/completions",
        completion_window="24h",
        metadata={"description": "nightly eval job"},
    )

    Checked against: Anthropic: Batch processing, OpenAI: Batch API guide

  6. Step 6Route easy requests to a cheaper model, behind an eval gate

    You end up with: A cheaper model handles the requests it can answer correctly, and the expensive model handles the rest.

    Model routing sends each request to the cheapest model that meets your quality bar. The simplest version needs no machine learning: pick the cheap model for whole features that are easy (classification, extraction, reformatting) and keep the stronger model for planning and open-ended reasoning. A second version tries the cheap model first and escalates when a check fails, for example when the output does not parse against a schema or when a confidence field is low.

    Do not switch on feel. Take the 30 to 50 real requests from your prerequisites, run them through both models and compare the outputs with whatever scoring you trust: exact match for labels, schema validity for extraction, a rubric for free text. Move a feature to the cheaper model only when the score holds. Then keep the gate: re-run the set whenever you change a prompt or a model.

    Count the escalations. If the cheap model fails a third of the time and every failure triggers a second call to the strong model, you pay for two calls on those requests and may spend more than before. Log which model answered, and watch the escalation rate next to cost.

    Cheap-first with an escalation check · python
    import json
    
    
    def classify(ticket_text, user_id):
        resp = call_claude(
            "ticket-classify", user_id,
            model="YOUR_CHEAP_MODEL_ID", max_tokens=100,
            messages=[{"role": "user", "content": "Return only JSON with keys category and urgency.\n\n" + ticket_text}],
        )
        try:
            data = json.loads(resp.content[0].text)
            if data.get("category") and data.get("urgency"):
                return data, "cheap"
        except (ValueError, IndexError):
            pass
        resp = call_claude(
            "ticket-classify-escalated", user_id,
            model="YOUR_STRONG_MODEL_ID", max_tokens=200,
            messages=[{"role": "user", "content": "Return only JSON with keys category and urgency.\n\n" + ticket_text}],
        )
        return json.loads(resp.content[0].text), "strong"

    Checked against: Anthropic: Messages API reference

  7. Step 7Cache whole answers and trim what you retrieve

    You end up with: Repeated questions skip the model entirely, and retrieval sends only the passages that matter.

    A response cache stores the final answer for an identical request. It suits deterministic features such as classification of a repeated input, FAQ-style questions and generated summaries of documents that have not changed. Key the cache on the model id, the full prompt and the settings, and give entries a lifetime so stale answers expire. Do not cache anything that depends on the user, or that contains personal data, unless the key includes the user.

    If you retrieve passages from a knowledge base, check how many you send. Many pipelines pass the top ten chunks when the top three contain the answer. Test the smaller number against your request set. Fewer, better chunks cost less and often answer better. Our cost of RAG page breaks down where retrieval spend goes, and how to build a RAG system covers choosing chunk size and count.

    Exact-match caching is the safe starting point. Semantic caching, which returns an earlier answer for a similar question, saves more calls but can return a wrong answer for a question that only looks similar. Add it only with a measured similarity threshold and a sample review.

    Exact-match response cache (standard library only) · python
    import hashlib
    import json
    import time
    
    _CACHE = {}  # replace with Redis or a database table in production
    TTL_SECONDS = 24 * 3600
    
    
    def cache_key(model, messages, **settings):
        blob = json.dumps({"m": model, "msgs": messages, "s": settings}, sort_keys=True)
        return hashlib.sha256(blob.encode("utf-8")).hexdigest()
    
    
    def cached_call(call_fn, model, messages, **settings):
        key = cache_key(model, messages, **settings)
        hit = _CACHE.get(key)
        if hit and time.time() - hit["at"] < TTL_SECONDS:
            return hit["value"]
        value = call_fn(model=model, messages=messages, **settings)
        _CACHE[key] = {"at": time.time(), "value": value}
        return value
  8. Step 8Set budgets and alerts, then price self-hosting honestly

    You end up with: Spend has a limit and an alert, and you know whether self-hosting could beat the API for your volume.

    Put a budget on each feature, and alert when daily spend passes a multiple of its trailing average. Runaway cost is almost always a loop, a retry storm or a prompt that grew. A gateway can enforce caps for you; how to set up an LLM gateway shows the steps.

    Self-hosting turns a variable cost into a fixed one, so it only wins at steady, high volume. The break-even is the monthly fixed cost divided by the saving per token. Fixed cost is the GPU rental (or purchase, spread over its life), power, and the engineering time to run and patch it. Use the throughput you measure in your own load test, at your latency target, with your prompts, because published benchmark numbers rarely match your workload.

    The worked calculation below uses hypothetical prices to show the method. Replace every number with your own before you decide anything. If the break-even volume is far above what you send today, stay on the API and revisit when volume grows. If you want to try it anyway, how to self-host an LLM is the next guide.

    Self-hosting break-even (hypothetical numbers, replace with yours)
    ItemFormulaExample (hypothetical)
    GPU rental per monthhourly price x 720 hours2.00 x 720 = 1,440
    Blended API priceyour logged mix of input and output5.00 per million tokens
    Break-even volumemonthly fixed cost / API price per million tokens1,440 / 5.00 = 288 million tokens a month
    Average rate neededmonthly tokens / seconds in a month (2,592,000)about 111 tokens a second, around the clock
    CheckCan one GPU serve that rate at your latency and quality?Measure it. If not, the break-even moves up.

Worked example: what each lever is worth (hypothetical numbers)

These figures are made up to show the arithmetic. They are not a price list, and your savings will differ. Assume an app sends 100,000 requests a month. Each has a 3,000-token static prefix, 200 tokens of changing input and 300 tokens of output. Hypothetical prices: 3.00 per million input tokens and 15.00 per million output tokens.

Baseline: input is 3,200 x 100,000 = 320 million tokens, which is 960. Output is 300 x 100,000 = 30 million tokens, which is 450. Total 1,410 a month.

Caching alone: suppose 95% of requests read the prefix from cache at 0.1 times the price and 5% write it at 1.25 times. The prefix costs 900 x (0.95 x 0.1 + 0.05 x 1.25) = 900 x 0.1575 = 141.75. Add 60 for the changing input and 450 for output: 651.75, a saving of about 54%.

Each lever applied alone to the same baseline (hypothetical)
LeverAssumptionMonthly costChange
BaselineNothing applied1,410-
Prompt caching95% of requests hit the cache651.75about -54%
Output capAverage output falls from 300 to 200 tokens1,260about -11%
Batch API40% of volume can wait, billed at 50%1,128-20%
Routing60% of requests move to a model at one fifth of the price733.20about -48%

Where cost cutting goes wrong

  • Switching to a cheaper model without an eval set. You save money and lose accuracy you cannot see.
  • Comparing list prices only. Reasoning tokens, retries, escalations and cache writes change the real cost per task. Our post on reasoning tokens and hidden cost explains the gap.
  • Caching prompts that change every call. You pay the write price and never get a read.
  • Truncating output with a low cap instead of asking for a shorter answer, then shipping cut-off JSON.
  • Treating a falling token price as a plan. Prices move, so build the measurement habit, not a one-off cut.

Troubleshooting

What you seeLikely causeFix
cache_read_input_tokens is always 0 on AnthropicThe prompt is shorter than the model minimum (512, 1,024, 2,048 or 4,096 tokens depending on the model), the prefix changes between calls, or more than 5 minutes passed on the default cache.Check the minimum for your model in the caching docs, move changing content (timestamps, ids) to the end, and use the 1-hour TTL if calls are further apart than 5 minutes.
input_tokens_details.cached_tokens is 0 on OpenAIThe prompt is below the minimum (1,024 tokens on the newest models), or the prefix differs between requests.Keep the static content first and byte-identical, add a stable prompt_cache_key for related requests, and read the caching guide for your model generation.
Answers end mid-sentence or the JSON does not parseThe output cap is lower than the answer needs. On Anthropic the stop_reason is max_tokens.Raise the cap from your logged 99th percentile, and ask for a shorter format in the prompt.
Batch results come back as expiredThe batch did not finish inside 24 hours, often because it was very large or demand was high.Resubmit only the expired requests, in smaller batches, and avoid sending a batch you need by a deadline.
Batch results show errored with an invalid_request_errorThe request body failed validation. Validation runs asynchronously, so errors appear in the results, not at submission.Fix the request body and send it again. Server errors can be retried as they are.
Cost went up after adding routingMany cheap-model failures trigger a second call to the strong model, so those requests cost more than before.Log the escalation rate. Narrow the cheap path to features where it passes your eval set, or improve the cheap prompt.
Your logged total does not match the invoiceCache writes, reasoning tokens, tool calls or batch discounts are missing from your calculation.Log every usage field from step 1 and price cache reads, cache writes and reasoning tokens separately.

Verify it worked

Next steps

Related guides

Frequently asked questions

What is the fastest way to reduce LLM costs?

Cap output length and turn on prompt caching for the static start of your prompts. Both are small code changes. Measure tokens per feature first so you know which calls to change, then re-measure after each change.

Does prompt caching really reduce cost?

Yes, when the start of the prompt repeats. As of 2026-10-06 Anthropic lists cache reads at 0.1 times the base input price for most models, and OpenAI lists 0.1 times on its newest models. Cache writes cost more, so a prompt that changes every call can cost more with caching than without.

Is a batch API cheaper?

Yes. Anthropic charges batch usage at 50% of standard prices and OpenAI states a 50% discount. The trade is time: results can take up to 24 hours, so use it for work nobody is waiting on.

How do I calculate LLM cost per request?

Multiply input tokens by the input price and output tokens by the output price, both per million tokens. Price cache reads, cache writes and reasoning tokens separately, using the usage fields the API returns. The token cost calculator does the sum for you.

What is model routing?

Model routing sends each request to the cheapest model that meets your quality bar, for example a small model for classification and a stronger one for planning. Test it against a set of real requests before you rely on it, and watch how often cheap answers get escalated.

When is self-hosting an LLM cheaper than an API?

When your volume is steady and high enough to keep the hardware busy. Divide your fixed monthly cost by your API price per million tokens to get the break-even volume, then check that your hardware can serve that rate at your latency target.

Do reasoning tokens cost money?

Yes. OpenAI reports them in output_tokens_details.reasoning_tokens and they count toward max_output_tokens. Log them, because they are a common reason a bill is higher than visible output suggests.

How Swfte can help

Everything above works with plain provider APIs. If you route traffic through Swfte Connect, the gateway records usage and cost per model, agent and workflow and can apply caps.

  • Usage and cost analytics: What the Swfte products measure and control today, and what is not available yet.
  • Swfte Connect: One OpenAI-compatible gateway with routing rules, caps and per-request records.
  • AI FinOps: The wider picture of AI spend control.

You can do every step above without Swfte. Per-step cost on every worker trace is not yet joined to the trace, and natural-language questions about spend are not available.

Missing a step or found a command that no longer works? Tell us, or request a how-to.

Sources and last verified

Commands, versions and facts in this guide were checked against the sources below on . Tools change quickly: if something differs from what you see, trust the official documentation and let us know.

  1. Anthropic: Prompt caching: cache_control syntax, 5-minute and 1-hour TTL, per-model minimum lengths, write and read multipliers, usage field names.
  2. OpenAI: Prompt caching: Caching on by default, 1,024-token minimum and 30-minute TTL on the newest models, prompt_cache_retention values, read multiplier, cached_tokens field.
  3. Anthropic: Batch processing: 50% batch pricing, 100,000 request and 256 MB limits, 24-hour expiry, 29-day result retention, Python create/retrieve/results code.
  4. OpenAI: Batch API guide: 50% discount, 24h completion window, JSONL fields, 50,000 request and 200 MB limits, Python upload and create code.
  5. Anthropic: Messages API reference: usage field names, max_tokens description, stop_reason values.
  6. OpenAI: Responses API reference: ResponseUsage fields (input_tokens, output_tokens, cached_tokens, reasoning_tokens) and max_output_tokens.
  7. PyPI: anthropic: pip install anthropic, Python 3.10 or later, client usage.
  8. PyPI: openai: pip install openai, Python 3.10 or later, Responses API usage.

Topics

  • cost
  • prompt caching
  • batch API
  • model routing
  • FinOps

Machine-readable copies: this guide as markdown, index of all guides (JSON). Canonical address: https://www.swfte.com/how-to-reduce-llm-costs.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.