Short answer
Measure tokens per feature, model and user before changing anything. Then work through the levers in order: cap output length, turn on prompt caching for the static start of every prompt, move work that can wait to a batch API at half price, route easy requests to a cheaper model behind an eval gate, and only then price self-hosting. Provider prices and cache rules change, so read each lever against the current docs.
The steps at a glance
- Log token usage on every model call
- Rank features by cost and find the top three
- Cap output length and tighten the prompt
- Turn on prompt caching for the static start of every prompt
- Move work that can wait to a batch API
- Route easy requests to a cheaper model, behind an eval gate
- Cache whole answers and trim what you retrieve
- Set budgets and alerts, then price self-hosting honestly
Before you start
Who this is for
- Developers and platform engineers whose monthly model bill has become a line item someone asks about.
- Technical leads who need a ranked list of savings with the arithmetic shown, not a list of tips.
- Teams running agents or retrieval pipelines, where a single run can make many model calls.
Probably not for you if
- People choosing their first model. Start with how to choose an LLM for your company.
- Teams that already know the number and only need a price list. Use the token cost calculator and the API pricing pages.
Prerequisites
- Access to the code that calls the model API, and permission to change it.
- Python 3.10 or later for the example code (the current Anthropic and OpenAI Python packages require it).
- An API key for at least one provider, kept in an environment variable and not in code.
- A set of 30 to 50 real requests with outputs you have judged acceptable. You need this before you try routing or a cheaper model.
- Time
- About 3 hours to add measurement and the first three levers; routing and batch work take longer
- Cost
- No new software to buy. Expect to spend a little on test calls while you compare models.
- Hardware
- None, unless you decide to self-host in the last step.
- Skill
- Comfortable reading and editing production code that calls an LLM API
Estimates are ours, not measurements, and move with your hardware, data and network.
Step 1Log token usage on every model call
You end up with: Every call writes one line saying which feature, model and user it belongs to, and how many tokens it used.
You cannot cut a bill you cannot attribute. Provider dashboards show spend per key, which is rarely the unit you care about. Wrap your model calls in one small function that writes a log line per call with a feature name, a model name, a user or tenant id and the token counts the provider returns.
The field names differ by provider. Anthropic returns
input_tokens,output_tokens,cache_creation_input_tokensandcache_read_input_tokensin theusageobject. OpenAI Responses returnsinput_tokens,output_tokens,input_tokens_details.cached_tokensandoutput_tokens_details.reasoning_tokens. Log all of them. Reasoning tokens and cache writes are billed, and they are the usual reason a bill does not match a naive count.Do not log prompt text here. Token counts and ids are enough for cost work, and keeping content out of the log avoids a second data-protection problem.
Install the SDKs · bash pip install anthropic openaiUsage logging wrapper (Anthropic and OpenAI Responses) · python import json import time from anthropic import Anthropic from openai import OpenAI anthropic_client = Anthropic() # reads ANTHROPIC_API_KEY openai_client = OpenAI() # reads OPENAI_API_KEY def log_usage(row): print(json.dumps(row)) # swap print for your logger or a file def call_claude(feature, user_id, **kwargs): start = time.time() resp = anthropic_client.messages.create(**kwargs) u = resp.usage log_usage({ "feature": feature, "user": user_id, "model": kwargs["model"], "input_tokens": u.input_tokens, "output_tokens": u.output_tokens, "cache_write_tokens": getattr(u, "cache_creation_input_tokens", 0) or 0, "cache_read_tokens": getattr(u, "cache_read_input_tokens", 0) or 0, "stop_reason": resp.stop_reason, "ms": int((time.time() - start) * 1000), }) return resp def call_openai(feature, user_id, **kwargs): start = time.time() resp = openai_client.responses.create(**kwargs) u = resp.usage log_usage({ "feature": feature, "user": user_id, "model": kwargs["model"], "input_tokens": u.input_tokens, "output_tokens": u.output_tokens, "cache_read_tokens": u.input_tokens_details.cached_tokens, "reasoning_tokens": u.output_tokens_details.reasoning_tokens, "ms": int((time.time() - start) * 1000), }) return respChecked against: Anthropic: Messages API reference, OpenAI: Responses API reference, PyPI: anthropic, PyPI: openai
Step 2Rank features by cost and find the top three
You end up with: You have a ranked list showing which three features, models or users cause most of the spend.
After a day or a week of logs, add the tokens up by feature and by model. Multiply by your current per-million-token prices to get money. Most teams find that two or three features cause most of the cost, and that within those features a few long prompts or very long outputs matter more than the average request.
Look for four patterns. A large static prefix repeated on every call (caching will help). Output far longer than the task needs (a cap will help). Work nobody is waiting for, such as nightly enrichment or evaluation runs (a batch API will help). And an expensive model doing simple work such as classification or extraction (routing will help). Write down which pattern each of your top three matches before you touch any code.
Total tokens by feature and model from the JSON-lines log · python import json import sys from collections import defaultdict totals = defaultdict(lambda: {"calls": 0, "in": 0, "out": 0, "cache_read": 0}) for line in open(sys.argv[1]): r = json.loads(line) t = totals[(r["feature"], r["model"])] t["calls"] += 1 t["in"] += r["input_tokens"] t["out"] += r["output_tokens"] t["cache_read"] += r.get("cache_read_tokens", 0) for (feature, model), t in sorted(totals.items(), key=lambda kv: -(kv[1]["in"] + kv[1]["out"])): print(f"{feature:24} {model:28} calls={t['calls']:7} in={t['in']:12} out={t['out']:10} cached={t['cache_read']:12}")Run it (save the script as rank_spend.py and your log as usage.jsonl) · bash python rank_spend.py usage.jsonlMatch each expensive pattern to a lever Pattern in your log Lever Step Large identical prefix, many calls Prompt caching 4 Output tokens much larger than the task needs Output cap and tighter prompt 3 Work nobody is waiting for Batch API 5 Frontier model on classification or extraction Model routing 6 Same question asked repeatedly Response cache 7 Step 3Cap output length and tighten the prompt
You end up with: Each feature has an explicit output limit and no longer asks for more text than the task needs.
Output tokens cost more than input tokens at every major provider, so unbounded output is the cheapest place to save. Set an explicit limit on every call. Anthropic calls it
max_tokens, the maximum number of tokens to generate before stopping. OpenAI Responses calls itmax_output_tokens, an upper bound that includes visible output and reasoning tokens.Set the cap from your log, not from a guess: take the 99th percentile of output tokens for the feature and add a margin. Then check the
stop_reasonfield on Anthropic responses. A value ofmax_tokensmeans the answer was cut off, so the cap is too low for that feature. Ask for the format you need in the prompt itself ("answer in two sentences", "return only the JSON object") rather than relying on the cap to truncate.Trim the input as well. Remove instructions the model no longer needs, drop examples that never change the answer, and stop re-sending a whole conversation when a short summary of it would do. Keep an eye on quality while you do this, because a shorter prompt that lowers accuracy is not a saving.
Checked against: Anthropic: Messages API reference, OpenAI: Responses API reference
Step 4Turn on prompt caching for the static start of every prompt
You end up with: Repeated prefixes are billed at the cheaper cache-read rate, and your log shows cache read tokens above zero.
Put everything that does not change first (system prompt, tool definitions, reference documents) and everything that changes last (the user message). Caching works on the unchanged beginning of the prompt, so one timestamp or request id near the top defeats it.
On Anthropic, add
cache_controlwith typeephemeralto the system block, or at the top level of the request for automatic caching. The default lifetime is 5 minutes, and"ttl": "1h"selects a 1-hour cache. As of 2026-10-06 the docs list a 5-minute cache write at 1.25 times the base input price, a 1-hour write at 2 times, and a cache read at 0.1 times for most models, with lower read multipliers on a few newer ones. Prompts shorter than the model minimum (512, 1,024, 2,048 or 4,096 tokens depending on the model) are not cached.On OpenAI, caching is on by default for supported models, so there is nothing to switch on. The docs give a 1,024-token minimum on the newest models, a 30-minute default lifetime there, and a read price of 0.1 times the uncached input price. Earlier models use
prompt_cache_retentionwithin_memoryor24h. A stableprompt_cache_keyhelps requests with a shared prefix land together. Check the hit rate in the field you logged: cached tokens divided by total input tokens.Anthropic: mark the static system prompt as cacheable · python LONG_STATIC_PROMPT = open("system_prompt.txt").read() # must exceed the model minimum resp = call_claude( "support-triage", "user-123", model="YOUR_MODEL_ID", max_tokens=400, system=[ { "type": "text", "text": LONG_STATIC_PROMPT, "cache_control": {"type": "ephemeral"}, } ], messages=[{"role": "user", "content": "Customer says the export button does nothing."}], )Anthropic: 1-hour cache for prompts used less often than every 5 minutes · json {"cache_control": {"type": "ephemeral", "ttl": "1h"}}Second call within the cache lifetime, from your log
{"feature": "support-triage", "cache_write_tokens": 0, "cache_read_tokens": <size of the cached prefix>, ...}Checked against: Anthropic: Prompt caching, OpenAI: Prompt caching
Step 5Move work that can wait to a batch API
You end up with: Offline jobs run through a batch endpoint at half the standard price.
Evaluation runs, nightly enrichment, bulk classification and back-fills do not need an answer in two seconds. Both Anthropic and OpenAI document a 50% discount for batch processing. On Anthropic, all batch usage is charged at 50% of standard prices. On OpenAI, the batch guide states a 50% discount and a completion window that can currently only be set to
24h.The limits matter when you plan the job. Anthropic allows up to 100,000 requests or 256 MB per batch, most batches finish within an hour, they expire after 24 hours, and results stay available for 29 days. OpenAI allows up to 50,000 requests and a 200 MB input file per batch. Each request carries a
custom_idso you can match results to inputs.The example below creates an Anthropic batch with two requests and then polls and reads the results. Streaming is not supported inside a batch. Check the provider docs for whether the batch discount combines with caching on your model, because Anthropic suggests the 1-hour cache for batches that share context.
Anthropic: create a batch · python import anthropic from anthropic.types.message_create_params import MessageCreateParamsNonStreaming from anthropic.types.messages.batch_create_params import Request client = anthropic.Anthropic() message_batch = client.messages.batches.create( requests=[ Request( custom_id="doc-0001", params=MessageCreateParamsNonStreaming( model="YOUR_MODEL_ID", max_tokens=300, messages=[{"role": "user", "content": "Classify this ticket: ..."}], ), ), Request( custom_id="doc-0002", params=MessageCreateParamsNonStreaming( model="YOUR_MODEL_ID", max_tokens=300, messages=[{"role": "user", "content": "Classify this ticket: ..."}], ), ), ] ) print(message_batch.id)Anthropic: poll until the batch has ended, then stream the results · python import time while True: batch = client.messages.batches.retrieve(message_batch.id) if batch.processing_status == "ended": break print("still processing...") time.sleep(60) for result in client.messages.batches.results(message_batch.id): outcome = result.result match outcome.type: case "succeeded": print("ok", result.custom_id) case "errored": print("errored", result.custom_id) case "expired": print("expired", result.custom_id)OpenAI: upload a JSONL file and create a batch · python from openai import OpenAI client = OpenAI() batch_input_file = client.files.create( file=open("batchinput.jsonl", "rb"), purpose="batch" ) batch = client.batches.create( input_file_id=batch_input_file.id, endpoint="/v1/chat/completions", completion_window="24h", metadata={"description": "nightly eval job"}, )Checked against: Anthropic: Batch processing, OpenAI: Batch API guide
Step 6Route easy requests to a cheaper model, behind an eval gate
You end up with: A cheaper model handles the requests it can answer correctly, and the expensive model handles the rest.
Model routing sends each request to the cheapest model that meets your quality bar. The simplest version needs no machine learning: pick the cheap model for whole features that are easy (classification, extraction, reformatting) and keep the stronger model for planning and open-ended reasoning. A second version tries the cheap model first and escalates when a check fails, for example when the output does not parse against a schema or when a confidence field is low.
Do not switch on feel. Take the 30 to 50 real requests from your prerequisites, run them through both models and compare the outputs with whatever scoring you trust: exact match for labels, schema validity for extraction, a rubric for free text. Move a feature to the cheaper model only when the score holds. Then keep the gate: re-run the set whenever you change a prompt or a model.
Count the escalations. If the cheap model fails a third of the time and every failure triggers a second call to the strong model, you pay for two calls on those requests and may spend more than before. Log which model answered, and watch the escalation rate next to cost.
Cheap-first with an escalation check · python import json def classify(ticket_text, user_id): resp = call_claude( "ticket-classify", user_id, model="YOUR_CHEAP_MODEL_ID", max_tokens=100, messages=[{"role": "user", "content": "Return only JSON with keys category and urgency.\n\n" + ticket_text}], ) try: data = json.loads(resp.content[0].text) if data.get("category") and data.get("urgency"): return data, "cheap" except (ValueError, IndexError): pass resp = call_claude( "ticket-classify-escalated", user_id, model="YOUR_STRONG_MODEL_ID", max_tokens=200, messages=[{"role": "user", "content": "Return only JSON with keys category and urgency.\n\n" + ticket_text}], ) return json.loads(resp.content[0].text), "strong"Checked against: Anthropic: Messages API reference
Step 7Cache whole answers and trim what you retrieve
You end up with: Repeated questions skip the model entirely, and retrieval sends only the passages that matter.
A response cache stores the final answer for an identical request. It suits deterministic features such as classification of a repeated input, FAQ-style questions and generated summaries of documents that have not changed. Key the cache on the model id, the full prompt and the settings, and give entries a lifetime so stale answers expire. Do not cache anything that depends on the user, or that contains personal data, unless the key includes the user.
If you retrieve passages from a knowledge base, check how many you send. Many pipelines pass the top ten chunks when the top three contain the answer. Test the smaller number against your request set. Fewer, better chunks cost less and often answer better. Our cost of RAG page breaks down where retrieval spend goes, and how to build a RAG system covers choosing chunk size and count.
Exact-match caching is the safe starting point. Semantic caching, which returns an earlier answer for a similar question, saves more calls but can return a wrong answer for a question that only looks similar. Add it only with a measured similarity threshold and a sample review.
Exact-match response cache (standard library only) · python import hashlib import json import time _CACHE = {} # replace with Redis or a database table in production TTL_SECONDS = 24 * 3600 def cache_key(model, messages, **settings): blob = json.dumps({"m": model, "msgs": messages, "s": settings}, sort_keys=True) return hashlib.sha256(blob.encode("utf-8")).hexdigest() def cached_call(call_fn, model, messages, **settings): key = cache_key(model, messages, **settings) hit = _CACHE.get(key) if hit and time.time() - hit["at"] < TTL_SECONDS: return hit["value"] value = call_fn(model=model, messages=messages, **settings) _CACHE[key] = {"at": time.time(), "value": value} return valueStep 8Set budgets and alerts, then price self-hosting honestly
You end up with: Spend has a limit and an alert, and you know whether self-hosting could beat the API for your volume.
Put a budget on each feature, and alert when daily spend passes a multiple of its trailing average. Runaway cost is almost always a loop, a retry storm or a prompt that grew. A gateway can enforce caps for you; how to set up an LLM gateway shows the steps.
Self-hosting turns a variable cost into a fixed one, so it only wins at steady, high volume. The break-even is the monthly fixed cost divided by the saving per token. Fixed cost is the GPU rental (or purchase, spread over its life), power, and the engineering time to run and patch it. Use the throughput you measure in your own load test, at your latency target, with your prompts, because published benchmark numbers rarely match your workload.
The worked calculation below uses hypothetical prices to show the method. Replace every number with your own before you decide anything. If the break-even volume is far above what you send today, stay on the API and revisit when volume grows. If you want to try it anyway, how to self-host an LLM is the next guide.
Self-hosting break-even (hypothetical numbers, replace with yours) Item Formula Example (hypothetical) GPU rental per month hourly price x 720 hours 2.00 x 720 = 1,440 Blended API price your logged mix of input and output 5.00 per million tokens Break-even volume monthly fixed cost / API price per million tokens 1,440 / 5.00 = 288 million tokens a month Average rate needed monthly tokens / seconds in a month (2,592,000) about 111 tokens a second, around the clock Check Can one GPU serve that rate at your latency and quality? Measure it. If not, the break-even moves up.
Worked example: what each lever is worth (hypothetical numbers)
These figures are made up to show the arithmetic. They are not a price list, and your savings will differ. Assume an app sends 100,000 requests a month. Each has a 3,000-token static prefix, 200 tokens of changing input and 300 tokens of output. Hypothetical prices: 3.00 per million input tokens and 15.00 per million output tokens.
Baseline: input is 3,200 x 100,000 = 320 million tokens, which is 960. Output is 300 x 100,000 = 30 million tokens, which is 450. Total 1,410 a month.
Caching alone: suppose 95% of requests read the prefix from cache at 0.1 times the price and 5% write it at 1.25 times. The prefix costs 900 x (0.95 x 0.1 + 0.05 x 1.25) = 900 x 0.1575 = 141.75. Add 60 for the changing input and 450 for output: 651.75, a saving of about 54%.
| Lever | Assumption | Monthly cost | Change |
|---|---|---|---|
| Baseline | Nothing applied | 1,410 | - |
| Prompt caching | 95% of requests hit the cache | 651.75 | about -54% |
| Output cap | Average output falls from 300 to 200 tokens | 1,260 | about -11% |
| Batch API | 40% of volume can wait, billed at 50% | 1,128 | -20% |
| Routing | 60% of requests move to a model at one fifth of the price | 733.20 | about -48% |
Where cost cutting goes wrong
- Switching to a cheaper model without an eval set. You save money and lose accuracy you cannot see.
- Comparing list prices only. Reasoning tokens, retries, escalations and cache writes change the real cost per task. Our post on reasoning tokens and hidden cost explains the gap.
- Caching prompts that change every call. You pay the write price and never get a read.
- Truncating output with a low cap instead of asking for a shorter answer, then shipping cut-off JSON.
- Treating a falling token price as a plan. Prices move, so build the measurement habit, not a one-off cut.
Troubleshooting
| What you see | Likely cause | Fix |
|---|---|---|
| cache_read_input_tokens is always 0 on Anthropic | The prompt is shorter than the model minimum (512, 1,024, 2,048 or 4,096 tokens depending on the model), the prefix changes between calls, or more than 5 minutes passed on the default cache. | Check the minimum for your model in the caching docs, move changing content (timestamps, ids) to the end, and use the 1-hour TTL if calls are further apart than 5 minutes. |
| input_tokens_details.cached_tokens is 0 on OpenAI | The prompt is below the minimum (1,024 tokens on the newest models), or the prefix differs between requests. | Keep the static content first and byte-identical, add a stable prompt_cache_key for related requests, and read the caching guide for your model generation. |
| Answers end mid-sentence or the JSON does not parse | The output cap is lower than the answer needs. On Anthropic the stop_reason is max_tokens. | Raise the cap from your logged 99th percentile, and ask for a shorter format in the prompt. |
Batch results come back as expired | The batch did not finish inside 24 hours, often because it was very large or demand was high. | Resubmit only the expired requests, in smaller batches, and avoid sending a batch you need by a deadline. |
Batch results show errored with an invalid_request_error | The request body failed validation. Validation runs asynchronously, so errors appear in the results, not at submission. | Fix the request body and send it again. Server errors can be retried as they are. |
| Cost went up after adding routing | Many cheap-model failures trigger a second call to the strong model, so those requests cost more than before. | Log the escalation rate. Narrow the cheap path to features where it passes your eval set, or improve the cheap prompt. |
| Your logged total does not match the invoice | Cache writes, reasoning tokens, tool calls or batch discounts are missing from your calculation. | Log every usage field from step 1 and price cache reads, cache writes and reasoning tokens separately. |
Verify it worked
Next steps
- Prompt caching explained: Deeper detail on how caching affects latency and price.
- Token cost calculator: Price your measured token mix across models.
- True cost per million tokens: See what the list price leaves out.
- How to set up an LLM gateway: Enforce caps, routing and logging in one place.
- How to monitor AI agents in production: Track cost per run next to errors and loops.
Related guides
- How to Set Up an LLM Gateway with LiteLLM (2026): Run the open-source LiteLLM proxy in Docker with a config file, add a model, issue virtual keys with budgets, set fallbacks, wire health checks, and harden it for production.
- How to Monitor AI Agents in Production (2026 Guide): Give every agent run an id, record each model and tool step as a span, redact before you store, alert on loops, tool failures and cost per run, and read a weekly sample by hand.
- How to Build a RAG System: Step-by-Step, Runs Locally: Build a retrieval-augmented generation system end to end on one machine, with hybrid search, per-group permissions, citations and a retrieval test set.
- How to Self-Host an LLM with vLLM (2026 Guide): Serve an open-weight model as a private, OpenAI-compatible endpoint on your own GPU server, with memory sizing, authentication, TLS, metrics and an upgrade routine.
- How to Choose an LLM for Your Company: A Scorecard: A selection process, not a leaderboard: requirements, a hosted and open-weight shortlist, a test on your own tasks, a weighted scorecard, licence and data-terms checks, and an exit plan.
Frequently asked questions
What is the fastest way to reduce LLM costs?
Cap output length and turn on prompt caching for the static start of your prompts. Both are small code changes. Measure tokens per feature first so you know which calls to change, then re-measure after each change.
Does prompt caching really reduce cost?
Yes, when the start of the prompt repeats. As of 2026-10-06 Anthropic lists cache reads at 0.1 times the base input price for most models, and OpenAI lists 0.1 times on its newest models. Cache writes cost more, so a prompt that changes every call can cost more with caching than without.
Is a batch API cheaper?
Yes. Anthropic charges batch usage at 50% of standard prices and OpenAI states a 50% discount. The trade is time: results can take up to 24 hours, so use it for work nobody is waiting on.
How do I calculate LLM cost per request?
Multiply input tokens by the input price and output tokens by the output price, both per million tokens. Price cache reads, cache writes and reasoning tokens separately, using the usage fields the API returns. The token cost calculator does the sum for you.
What is model routing?
Model routing sends each request to the cheapest model that meets your quality bar, for example a small model for classification and a stronger one for planning. Test it against a set of real requests before you rely on it, and watch how often cheap answers get escalated.
When is self-hosting an LLM cheaper than an API?
When your volume is steady and high enough to keep the hardware busy. Divide your fixed monthly cost by your API price per million tokens to get the break-even volume, then check that your hardware can serve that rate at your latency target.
Do reasoning tokens cost money?
Yes. OpenAI reports them in output_tokens_details.reasoning_tokens and they count toward max_output_tokens. Log them, because they are a common reason a bill is higher than visible output suggests.
How Swfte can help
Everything above works with plain provider APIs. If you route traffic through Swfte Connect, the gateway records usage and cost per model, agent and workflow and can apply caps.
- Usage and cost analytics: What the Swfte products measure and control today, and what is not available yet.
- Swfte Connect: One OpenAI-compatible gateway with routing rules, caps and per-request records.
- AI FinOps: The wider picture of AI spend control.
You can do every step above without Swfte. Per-step cost on every worker trace is not yet joined to the trace, and natural-language questions about spend are not available.
Missing a step or found a command that no longer works? Tell us, or request a how-to.
Sources and last verified
Commands, versions and facts in this guide were checked against the sources below on . Tools change quickly: if something differs from what you see, trust the official documentation and let us know.
- Anthropic: Prompt caching: cache_control syntax, 5-minute and 1-hour TTL, per-model minimum lengths, write and read multipliers, usage field names.
- OpenAI: Prompt caching: Caching on by default, 1,024-token minimum and 30-minute TTL on the newest models, prompt_cache_retention values, read multiplier, cached_tokens field.
- Anthropic: Batch processing: 50% batch pricing, 100,000 request and 256 MB limits, 24-hour expiry, 29-day result retention, Python create/retrieve/results code.
- OpenAI: Batch API guide: 50% discount, 24h completion window, JSONL fields, 50,000 request and 200 MB limits, Python upload and create code.
- Anthropic: Messages API reference: usage field names, max_tokens description, stop_reason values.
- OpenAI: Responses API reference: ResponseUsage fields (input_tokens, output_tokens, cached_tokens, reasoning_tokens) and max_output_tokens.
- PyPI: anthropic: pip install anthropic, Python 3.10 or later, client usage.
- PyPI: openai: pip install openai, Python 3.10 or later, Responses API usage.
Topics
- cost
- prompt caching
- batch API
- model routing
- FinOps
Machine-readable copies: this guide as markdown, index of all guides (JSON). Canonical address: https://www.swfte.com/how-to-reduce-llm-costs.