# How to reduce LLM costs

Canonical: https://www.swfte.com/how-to-reduce-llm-costs
Last verified: 2026-10-06
Difficulty: Intermediate
Time: About 3 hours to add measurement and the first three levers; routing and batch work take longer
Cost: No new software to buy. Expect to spend a little on test calls while you compare models.
Hardware: None, unless you decide to self-host in the last step.

## Short answer

Measure tokens per feature, model and user before changing anything. Then work through the levers in order: cap output length, turn on prompt caching for the static start of every prompt, move work that can wait to a batch API at half price, route easy requests to a cheaper model behind an eval gate, and only then price self-hosting. Provider prices and cache rules change, so read each lever against the current docs.

## Who this is for

- Developers and platform engineers whose monthly model bill has become a line item someone asks about.
- Technical leads who need a ranked list of savings with the arithmetic shown, not a list of tips.
- Teams running agents or retrieval pipelines, where a single run can make many model calls.

Not for:
- People choosing their first model. Start with [how to choose an LLM for your company](https://www.swfte.com/how-to-choose-an-llm-for-your-company).
- Teams that already know the number and only need a price list. Use the [token cost calculator](https://www.swfte.com/ai/token-cost-calculator) and the [API pricing pages](https://www.swfte.com/api-pricing).

## Prerequisites

- Access to the code that calls the model API, and permission to change it.
- Python 3.10 or later for the example code (the current Anthropic and OpenAI Python packages require it).
- An API key for at least one provider, kept in an environment variable and not in code.
- A set of 30 to 50 real requests with outputs you have judged acceptable. You need this before you try routing or a cheaper model.

## Steps

### Step 1: Log token usage on every model call

Outcome: Every call writes one line saying which feature, model and user it belongs to, and how many tokens it used.

You cannot cut a bill you cannot attribute. Provider dashboards show spend per key, which is rarely the unit you care about. Wrap your model calls in one small function that writes a log line per call with a feature name, a model name, a user or tenant id and the token counts the provider returns.

The field names differ by provider. Anthropic returns `input_tokens`, `output_tokens`, `cache_creation_input_tokens` and `cache_read_input_tokens` in the `usage` object. OpenAI Responses returns `input_tokens`, `output_tokens`, `input_tokens_details.cached_tokens` and `output_tokens_details.reasoning_tokens`. Log all of them. Reasoning tokens and cache writes are billed, and they are the usual reason a bill does not match a naive count.

Do not log prompt text here. Token counts and ids are enough for cost work, and keeping content out of the log avoids a second data-protection problem.

Install the SDKs:

```bash
pip install anthropic openai
```

Usage logging wrapper (Anthropic and OpenAI Responses):

```python
import json
import time

from anthropic import Anthropic
from openai import OpenAI

anthropic_client = Anthropic()  # reads ANTHROPIC_API_KEY
openai_client = OpenAI()        # reads OPENAI_API_KEY


def log_usage(row):
    print(json.dumps(row))  # swap print for your logger or a file


def call_claude(feature, user_id, **kwargs):
    start = time.time()
    resp = anthropic_client.messages.create(**kwargs)
    u = resp.usage
    log_usage({
        "feature": feature,
        "user": user_id,
        "model": kwargs["model"],
        "input_tokens": u.input_tokens,
        "output_tokens": u.output_tokens,
        "cache_write_tokens": getattr(u, "cache_creation_input_tokens", 0) or 0,
        "cache_read_tokens": getattr(u, "cache_read_input_tokens", 0) or 0,
        "stop_reason": resp.stop_reason,
        "ms": int((time.time() - start) * 1000),
    })
    return resp


def call_openai(feature, user_id, **kwargs):
    start = time.time()
    resp = openai_client.responses.create(**kwargs)
    u = resp.usage
    log_usage({
        "feature": feature,
        "user": user_id,
        "model": kwargs["model"],
        "input_tokens": u.input_tokens,
        "output_tokens": u.output_tokens,
        "cache_read_tokens": u.input_tokens_details.cached_tokens,
        "reasoning_tokens": u.output_tokens_details.reasoning_tokens,
        "ms": int((time.time() - start) * 1000),
    })
    return resp
```

> NOTE: Use the model ids from your own provider account. The examples in the official docs change with each release, so this guide leaves the model argument to you.

### Step 2: Rank features by cost and find the top three

Outcome: You have a ranked list showing which three features, models or users cause most of the spend.

After a day or a week of logs, add the tokens up by feature and by model. Multiply by your current per-million-token prices to get money. Most teams find that two or three features cause most of the cost, and that within those features a few long prompts or very long outputs matter more than the average request.

Look for four patterns. A large static prefix repeated on every call (caching will help). Output far longer than the task needs (a cap will help). Work nobody is waiting for, such as nightly enrichment or evaluation runs (a batch API will help). And an expensive model doing simple work such as classification or extraction (routing will help). Write down which pattern each of your top three matches before you touch any code.

Total tokens by feature and model from the JSON-lines log:

```python
import json
import sys
from collections import defaultdict

totals = defaultdict(lambda: {"calls": 0, "in": 0, "out": 0, "cache_read": 0})
for line in open(sys.argv[1]):
    r = json.loads(line)
    t = totals[(r["feature"], r["model"])]
    t["calls"] += 1
    t["in"] += r["input_tokens"]
    t["out"] += r["output_tokens"]
    t["cache_read"] += r.get("cache_read_tokens", 0)

for (feature, model), t in sorted(totals.items(), key=lambda kv: -(kv[1]["in"] + kv[1]["out"])):
    print(f"{feature:24} {model:28} calls={t['calls']:7} in={t['in']:12} out={t['out']:10} cached={t['cache_read']:12}")
```

Run it (save the script as rank_spend.py and your log as usage.jsonl):

```bash
python rank_spend.py usage.jsonl
```

**Match each expensive pattern to a lever**

| Pattern in your log | Lever | Step |
| --- | --- | --- |
| Large identical prefix, many calls | Prompt caching | 4 |
| Output tokens much larger than the task needs | Output cap and tighter prompt | 3 |
| Work nobody is waiting for | Batch API | 5 |
| Frontier model on classification or extraction | Model routing | 6 |
| Same question asked repeatedly | Response cache | 7 |

### Step 3: Cap output length and tighten the prompt

Outcome: Each feature has an explicit output limit and no longer asks for more text than the task needs.

Output tokens cost more than input tokens at every major provider, so unbounded output is the cheapest place to save. Set an explicit limit on every call. Anthropic calls it `max_tokens`, the maximum number of tokens to generate before stopping. OpenAI Responses calls it `max_output_tokens`, an upper bound that includes visible output and reasoning tokens.

Set the cap from your log, not from a guess: take the 99th percentile of output tokens for the feature and add a margin. Then check the `stop_reason` field on Anthropic responses. A value of `max_tokens` means the answer was cut off, so the cap is too low for that feature. Ask for the format you need in the prompt itself ("answer in two sentences", "return only the JSON object") rather than relying on the cap to truncate.

Trim the input as well. Remove instructions the model no longer needs, drop examples that never change the answer, and stop re-sending a whole conversation when a short summary of it would do. Keep an eye on quality while you do this, because a shorter prompt that lowers accuracy is not a saving.

> WARNING: With reasoning models, a low `max_output_tokens` can use the whole budget on hidden reasoning and return no visible answer. Raise the cap in steps and watch the reasoning-token count you logged in step 1.

### Step 4: Turn on prompt caching for the static start of every prompt

Outcome: Repeated prefixes are billed at the cheaper cache-read rate, and your log shows cache read tokens above zero.

Put everything that does not change first (system prompt, tool definitions, reference documents) and everything that changes last (the user message). Caching works on the unchanged beginning of the prompt, so one timestamp or request id near the top defeats it.

On Anthropic, add `cache_control` with type `ephemeral` to the system block, or at the top level of the request for automatic caching. The default lifetime is 5 minutes, and `"ttl": "1h"` selects a 1-hour cache. As of 2026-10-06 the docs list a 5-minute cache write at 1.25 times the base input price, a 1-hour write at 2 times, and a cache read at 0.1 times for most models, with lower read multipliers on a few newer ones. Prompts shorter than the model minimum (512, 1,024, 2,048 or 4,096 tokens depending on the model) are not cached.

On OpenAI, caching is on by default for supported models, so there is nothing to switch on. The docs give a 1,024-token minimum on the newest models, a 30-minute default lifetime there, and a read price of 0.1 times the uncached input price. Earlier models use `prompt_cache_retention` with `in_memory` or `24h`. A stable `prompt_cache_key` helps requests with a shared prefix land together. Check the hit rate in the field you logged: cached tokens divided by total input tokens.

Anthropic: mark the static system prompt as cacheable:

```python
LONG_STATIC_PROMPT = open("system_prompt.txt").read()  # must exceed the model minimum

resp = call_claude(
    "support-triage",
    "user-123",
    model="YOUR_MODEL_ID",
    max_tokens=400,
    system=[
        {
            "type": "text",
            "text": LONG_STATIC_PROMPT,
            "cache_control": {"type": "ephemeral"},
        }
    ],
    messages=[{"role": "user", "content": "Customer says the export button does nothing."}],
)
```

Anthropic: 1-hour cache for prompts used less often than every 5 minutes:

```json
{"cache_control": {"type": "ephemeral", "ttl": "1h"}}
```

Second call within the cache lifetime, from your log:

```text
{"feature": "support-triage", "cache_write_tokens": 0, "cache_read_tokens": <size of the cached prefix>, ...}
```

> TIP: Run the same request twice, a few seconds apart. The first call should show cache write tokens and the second should show cache read tokens. If both are zero, go to the troubleshooting table.

### Step 5: Move work that can wait to a batch API

Outcome: Offline jobs run through a batch endpoint at half the standard price.

Evaluation runs, nightly enrichment, bulk classification and back-fills do not need an answer in two seconds. Both Anthropic and OpenAI document a 50% discount for batch processing. On Anthropic, all batch usage is charged at 50% of standard prices. On OpenAI, the batch guide states a 50% discount and a completion window that can currently only be set to `24h`.

The limits matter when you plan the job. Anthropic allows up to 100,000 requests or 256 MB per batch, most batches finish within an hour, they expire after 24 hours, and results stay available for 29 days. OpenAI allows up to 50,000 requests and a 200 MB input file per batch. Each request carries a `custom_id` so you can match results to inputs.

The example below creates an Anthropic batch with two requests and then polls and reads the results. Streaming is not supported inside a batch. Check the provider docs for whether the batch discount combines with caching on your model, because Anthropic suggests the 1-hour cache for batches that share context.

Anthropic: create a batch:

```python
import anthropic
from anthropic.types.message_create_params import MessageCreateParamsNonStreaming
from anthropic.types.messages.batch_create_params import Request

client = anthropic.Anthropic()

message_batch = client.messages.batches.create(
    requests=[
        Request(
            custom_id="doc-0001",
            params=MessageCreateParamsNonStreaming(
                model="YOUR_MODEL_ID",
                max_tokens=300,
                messages=[{"role": "user", "content": "Classify this ticket: ..."}],
            ),
        ),
        Request(
            custom_id="doc-0002",
            params=MessageCreateParamsNonStreaming(
                model="YOUR_MODEL_ID",
                max_tokens=300,
                messages=[{"role": "user", "content": "Classify this ticket: ..."}],
            ),
        ),
    ]
)
print(message_batch.id)
```

Anthropic: poll until the batch has ended, then stream the results:

```python
import time

while True:
    batch = client.messages.batches.retrieve(message_batch.id)
    if batch.processing_status == "ended":
        break
    print("still processing...")
    time.sleep(60)

for result in client.messages.batches.results(message_batch.id):
    outcome = result.result
    match outcome.type:
        case "succeeded":
            print("ok", result.custom_id)
        case "errored":
            print("errored", result.custom_id)
        case "expired":
            print("expired", result.custom_id)
```

OpenAI: upload a JSONL file and create a batch:

```python
from openai import OpenAI

client = OpenAI()

batch_input_file = client.files.create(
    file=open("batchinput.jsonl", "rb"), purpose="batch"
)

batch = client.batches.create(
    input_file_id=batch_input_file.id,
    endpoint="/v1/chat/completions",
    completion_window="24h",
    metadata={"description": "nightly eval job"},
)
```

### Step 6: Route easy requests to a cheaper model, behind an eval gate

Outcome: A cheaper model handles the requests it can answer correctly, and the expensive model handles the rest.

Model routing sends each request to the cheapest model that meets your quality bar. The simplest version needs no machine learning: pick the cheap model for whole features that are easy (classification, extraction, reformatting) and keep the stronger model for planning and open-ended reasoning. A second version tries the cheap model first and escalates when a check fails, for example when the output does not parse against a schema or when a confidence field is low.

Do not switch on feel. Take the 30 to 50 real requests from your prerequisites, run them through both models and compare the outputs with whatever scoring you trust: exact match for labels, schema validity for extraction, a rubric for free text. Move a feature to the cheaper model only when the score holds. Then keep the gate: re-run the set whenever you change a prompt or a model.

Count the escalations. If the cheap model fails a third of the time and every failure triggers a second call to the strong model, you pay for two calls on those requests and may spend more than before. Log which model answered, and watch the escalation rate next to cost.

Cheap-first with an escalation check:

```python
import json


def classify(ticket_text, user_id):
    resp = call_claude(
        "ticket-classify", user_id,
        model="YOUR_CHEAP_MODEL_ID", max_tokens=100,
        messages=[{"role": "user", "content": "Return only JSON with keys category and urgency.\n\n" + ticket_text}],
    )
    try:
        data = json.loads(resp.content[0].text)
        if data.get("category") and data.get("urgency"):
            return data, "cheap"
    except (ValueError, IndexError):
        pass
    resp = call_claude(
        "ticket-classify-escalated", user_id,
        model="YOUR_STRONG_MODEL_ID", max_tokens=200,
        messages=[{"role": "user", "content": "Return only JSON with keys category and urgency.\n\n" + ticket_text}],
    )
    return json.loads(resp.content[0].text), "strong"
```

> TIP: Our [model mixing guide](https://www.swfte.com/ai/model-mixing-cost-savings) and the post on [routing for cost](https://www.swfte.com/blog/ai-model-routing-cost-optimization-2025) go deeper on routing strategies.

### Step 7: Cache whole answers and trim what you retrieve

Outcome: Repeated questions skip the model entirely, and retrieval sends only the passages that matter.

A response cache stores the final answer for an identical request. It suits deterministic features such as classification of a repeated input, FAQ-style questions and generated summaries of documents that have not changed. Key the cache on the model id, the full prompt and the settings, and give entries a lifetime so stale answers expire. Do not cache anything that depends on the user, or that contains personal data, unless the key includes the user.

If you retrieve passages from a knowledge base, check how many you send. Many pipelines pass the top ten chunks when the top three contain the answer. Test the smaller number against your request set. Fewer, better chunks cost less and often answer better. Our [cost of RAG](https://www.swfte.com/cost-of-rag) page breaks down where retrieval spend goes, and [how to build a RAG system](https://www.swfte.com/how-to-build-a-rag-system) covers choosing chunk size and count.

Exact-match caching is the safe starting point. Semantic caching, which returns an earlier answer for a similar question, saves more calls but can return a wrong answer for a question that only looks similar. Add it only with a measured similarity threshold and a sample review.

Exact-match response cache (standard library only):

```python
import hashlib
import json
import time

_CACHE = {}  # replace with Redis or a database table in production
TTL_SECONDS = 24 * 3600


def cache_key(model, messages, **settings):
    blob = json.dumps({"m": model, "msgs": messages, "s": settings}, sort_keys=True)
    return hashlib.sha256(blob.encode("utf-8")).hexdigest()


def cached_call(call_fn, model, messages, **settings):
    key = cache_key(model, messages, **settings)
    hit = _CACHE.get(key)
    if hit and time.time() - hit["at"] < TTL_SECONDS:
        return hit["value"]
    value = call_fn(model=model, messages=messages, **settings)
    _CACHE[key] = {"at": time.time(), "value": value}
    return value
```

### Step 8: Set budgets and alerts, then price self-hosting honestly

Outcome: Spend has a limit and an alert, and you know whether self-hosting could beat the API for your volume.

Put a budget on each feature, and alert when daily spend passes a multiple of its trailing average. Runaway cost is almost always a loop, a retry storm or a prompt that grew. A gateway can enforce caps for you; [how to set up an LLM gateway](https://www.swfte.com/how-to-set-up-an-llm-gateway) shows the steps.

Self-hosting turns a variable cost into a fixed one, so it only wins at steady, high volume. The break-even is the monthly fixed cost divided by the saving per token. Fixed cost is the GPU rental (or purchase, spread over its life), power, and the engineering time to run and patch it. Use the throughput you measure in your own load test, at your latency target, with your prompts, because published benchmark numbers rarely match your workload.

The worked calculation below uses hypothetical prices to show the method. Replace every number with your own before you decide anything. If the break-even volume is far above what you send today, stay on the API and revisit when volume grows. If you want to try it anyway, [how to self-host an LLM](https://www.swfte.com/how-to-self-host-an-llm) is the next guide.

**Self-hosting break-even (hypothetical numbers, replace with yours)**

| Item | Formula | Example (hypothetical) |
| --- | --- | --- |
| GPU rental per month | hourly price x 720 hours | 2.00 x 720 = 1,440 |
| Blended API price | your logged mix of input and output | 5.00 per million tokens |
| Break-even volume | monthly fixed cost / API price per million tokens | 1,440 / 5.00 = 288 million tokens a month |
| Average rate needed | monthly tokens / seconds in a month (2,592,000) | about 111 tokens a second, around the clock |
| Check | Can one GPU serve that rate at your latency and quality? | Measure it. If not, the break-even moves up. |

## Worked example: what each lever is worth (hypothetical numbers)

These figures are made up to show the arithmetic. They are not a price list, and your savings will differ. Assume an app sends 100,000 requests a month. Each has a 3,000-token static prefix, 200 tokens of changing input and 300 tokens of output. Hypothetical prices: 3.00 per million input tokens and 15.00 per million output tokens.

Baseline: input is 3,200 x 100,000 = 320 million tokens, which is 960. Output is 300 x 100,000 = 30 million tokens, which is 450. Total 1,410 a month.

Caching alone: suppose 95% of requests read the prefix from cache at 0.1 times the price and 5% write it at 1.25 times. The prefix costs 900 x (0.95 x 0.1 + 0.05 x 1.25) = 900 x 0.1575 = 141.75. Add 60 for the changing input and 450 for output: 651.75, a saving of about 54%.

**Each lever applied alone to the same baseline (hypothetical)**

| Lever | Assumption | Monthly cost | Change |
| --- | --- | --- | --- |
| Baseline | Nothing applied | 1,410 | - |
| Prompt caching | 95% of requests hit the cache | 651.75 | about -54% |
| Output cap | Average output falls from 300 to 200 tokens | 1,260 | about -11% |
| Batch API | 40% of volume can wait, billed at 50% | 1,128 | -20% |
| Routing | 60% of requests move to a model at one fifth of the price | 733.20 | about -48% |

> NOTE: The levers do not simply add up, because each acts on part of the same cost. Apply them one at a time, re-measure after each, and stop when the next saving is smaller than the engineering time it costs.

## Where cost cutting goes wrong

- Switching to a cheaper model without an eval set. You save money and lose accuracy you cannot see.
- Comparing list prices only. Reasoning tokens, retries, escalations and cache writes change the real cost per task. Our post on [reasoning tokens and hidden cost](https://www.swfte.com/blog/reasoning-token-hidden-cost-llm-pricing-2026) explains the gap.
- Caching prompts that change every call. You pay the write price and never get a read.
- Truncating output with a low cap instead of asking for a shorter answer, then shipping cut-off JSON.
- Treating a falling token price as a plan. Prices move, so build the measurement habit, not a one-off cut.

## Troubleshooting

| Symptom | Likely cause | Fix |
| --- | --- | --- |
| cache_read_input_tokens is always 0 on Anthropic | The prompt is shorter than the model minimum (512, 1,024, 2,048 or 4,096 tokens depending on the model), the prefix changes between calls, or more than 5 minutes passed on the default cache. | Check the minimum for your model in the caching docs, move changing content (timestamps, ids) to the end, and use the 1-hour TTL if calls are further apart than 5 minutes. |
| input_tokens_details.cached_tokens is 0 on OpenAI | The prompt is below the minimum (1,024 tokens on the newest models), or the prefix differs between requests. | Keep the static content first and byte-identical, add a stable `prompt_cache_key` for related requests, and read the caching guide for your model generation. |
| Answers end mid-sentence or the JSON does not parse | The output cap is lower than the answer needs. On Anthropic the `stop_reason` is `max_tokens`. | Raise the cap from your logged 99th percentile, and ask for a shorter format in the prompt. |
| Batch results come back as `expired` | The batch did not finish inside 24 hours, often because it was very large or demand was high. | Resubmit only the expired requests, in smaller batches, and avoid sending a batch you need by a deadline. |
| Batch results show `errored` with an `invalid_request_error` | The request body failed validation. Validation runs asynchronously, so errors appear in the results, not at submission. | Fix the request body and send it again. Server errors can be retried as they are. |
| Cost went up after adding routing | Many cheap-model failures trigger a second call to the strong model, so those requests cost more than before. | Log the escalation rate. Narrow the cheap path to features where it passes your eval set, or improve the cheap prompt. |
| Your logged total does not match the invoice | Cache writes, reasoning tokens, tool calls or batch discounts are missing from your calculation. | Log every usage field from step 1 and price cache reads, cache writes and reasoning tokens separately. |

## Verify it worked

- [ ] Every model call writes a log line with feature, model, user, input tokens, output tokens and cache tokens.
- [ ] You can name the three features that cause most of the spend, with their share.
- [ ] Every call has an explicit output cap, and the share of responses that stop at the cap is near zero.
- [ ] A repeated request shows cache read tokens above zero, and the cache hit rate is on a dashboard.
- [ ] At least one offline job runs through a batch API, and its results are matched back by `custom_id`.
- [ ] Any cheaper model in use passed your 30 to 50 request eval set, and the escalation rate is logged.
- [ ] A budget and an alert exist per feature, and you have a written break-even figure for self-hosting.

## Next steps

- [Prompt caching explained](https://www.swfte.com/prompt-caching): Deeper detail on how caching affects latency and price.
- [Token cost calculator](https://www.swfte.com/ai/token-cost-calculator): Price your measured token mix across models.
- [True cost per million tokens](https://www.swfte.com/ai/per-million-tokens-true-cost): See what the list price leaves out.
- [How to set up an LLM gateway](https://www.swfte.com/how-to-set-up-an-llm-gateway): Enforce caps, routing and logging in one place.
- [How to monitor AI agents in production](https://www.swfte.com/how-to-monitor-ai-agents-in-production): Track cost per run next to errors and loops.

## FAQ

### What is the fastest way to reduce LLM costs?

Cap output length and turn on prompt caching for the static start of your prompts. Both are small code changes. Measure tokens per feature first so you know which calls to change, then re-measure after each change.

### Does prompt caching really reduce cost?

Yes, when the start of the prompt repeats. As of 2026-10-06 Anthropic lists cache reads at 0.1 times the base input price for most models, and OpenAI lists 0.1 times on its newest models. Cache writes cost more, so a prompt that changes every call can cost more with caching than without.

### Is a batch API cheaper?

Yes. Anthropic charges batch usage at 50% of standard prices and OpenAI states a 50% discount. The trade is time: results can take up to 24 hours, so use it for work nobody is waiting on.

### How do I calculate LLM cost per request?

Multiply input tokens by the input price and output tokens by the output price, both per million tokens. Price cache reads, cache writes and reasoning tokens separately, using the usage fields the API returns. The [token cost calculator](https://www.swfte.com/ai/token-cost-calculator) does the sum for you.

### What is model routing?

Model routing sends each request to the cheapest model that meets your quality bar, for example a small model for classification and a stronger one for planning. Test it against a set of real requests before you rely on it, and watch how often cheap answers get escalated.

### When is self-hosting an LLM cheaper than an API?

When your volume is steady and high enough to keep the hardware busy. Divide your fixed monthly cost by your API price per million tokens to get the break-even volume, then check that your hardware can serve that rate at your latency target.

### Do reasoning tokens cost money?

Yes. OpenAI reports them in `output_tokens_details.reasoning_tokens` and they count toward `max_output_tokens`. Log them, because they are a common reason a bill is higher than visible output suggests.

## How Swfte can help

Everything above works with plain provider APIs. If you route traffic through Swfte Connect, the gateway records usage and cost per model, agent and workflow and can apply caps.

- [Usage and cost analytics](https://www.swfte.com/platform/intelligence/usage-and-cost-analytics): What the Swfte products measure and control today, and what is not available yet.
- [Swfte Connect](https://www.swfte.com/products/connect): One OpenAI-compatible gateway with routing rules, caps and per-request records.
- [AI FinOps](https://www.swfte.com/ai-finops): The wider picture of AI spend control.

You can do every step above without Swfte. Per-step cost on every worker trace is not yet joined to the trace, and natural-language questions about spend are not available.

## Sources

- [Anthropic: Prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching): cache_control syntax, 5-minute and 1-hour TTL, per-model minimum lengths, write and read multipliers, usage field names.
- [OpenAI: Prompt caching](https://developers.openai.com/api/docs/guides/prompt-caching): Caching on by default, 1,024-token minimum and 30-minute TTL on the newest models, prompt_cache_retention values, read multiplier, cached_tokens field.
- [Anthropic: Batch processing](https://platform.claude.com/docs/en/build-with-claude/batch-processing): 50% batch pricing, 100,000 request and 256 MB limits, 24-hour expiry, 29-day result retention, Python create/retrieve/results code.
- [OpenAI: Batch API guide](https://developers.openai.com/api/docs/guides/batch): 50% discount, 24h completion window, JSONL fields, 50,000 request and 200 MB limits, Python upload and create code.
- [Anthropic: Messages API reference](https://platform.claude.com/docs/en/api/messages/create): usage field names, max_tokens description, stop_reason values.
- [OpenAI: Responses API reference](https://developers.openai.com/api/reference/resources/responses/methods/create): ResponseUsage fields (input_tokens, output_tokens, cached_tokens, reasoning_tokens) and max_output_tokens.
- [PyPI: anthropic](https://pypi.org/project/anthropic/): pip install anthropic, Python 3.10 or later, client usage.
- [PyPI: openai](https://pypi.org/project/openai/): pip install openai, Python 3.10 or later, Responses API usage.

Last verified against these sources on 2026-10-06.
