Claude Opus 5.5 Pricing: 20% Cheaper Per Token, Anthropic Says 40% Per Task
Opus 5.5 is 20% cheaper per token and Anthropic estimates 40% per task. Sonnet 5.5 holds $2/$10 and wins on tokens.
Picture an engineering lead whose Claude bill has become a line item people ask about in planning meetings. Her coding agents run on Opus 5. On 22 September Anthropic released Claude Opus 5.5 at $4 and $20 per million tokens, 20% below Opus 5's $5 and $25. On 28 September it released Claude Sonnet 5.5 at $2 and $10, the same rate card as Sonnet 5. Her first question is the obvious one. How far will the invoice fall?
The honest answer has two parts, and the announcements blur them. Opus 5.5 is 20% cheaper per token, and 60% cheaper on cache reads. Anthropic also says it uses fewer tokens per task, and it estimates the two together "net out to a 40% drop in costs" on typical workloads at default settings. Sonnet 5.5 is not cheaper per token at all. Anthropic says it is cheaper per task, by up to 30%, because it needs fewer tokens and fewer steps.
This post keeps those two kinds of cost apart and works an example through with numbers you can replace with your own. Every Anthropic benchmark figure is vendor-run.
What shipped, and what it costs
Both models are closed, both have a 1M-token context window, and both allow up to 128K tokens of output in a synchronous call. Anthropic's pricing page says Claude 4.6 and later models include the full 1M window at standard pricing, so neither carries a long-context surcharge. Here is the rate card, per million tokens.
| Opus 5.5 | Opus 5 | Sonnet 5.5 | Fable 5.1 | |
|---|---|---|---|---|
| Input | $4 | $5 | $2 | $10 |
| Output | $20 | $25 | $10 | $50 |
| Cache read | $0.20 (0.05x) | $0.50 | $0.20 (0.1x) | $0.25 (0.025x) |
Against Opus 5, input, output and cache writes are each 20% lower on Opus 5.5 ($5 against $6.25 for a five-minute write). Cache reads fall from $0.50 to $0.20, which is 60% lower. Against Fable 5.1, Opus 5.5 is 60% cheaper on input and output. Those percentages are my arithmetic on the table, not Anthropic's claims.
Look at the cache-read row. Opus 5.5 reads cache at half the usual 0.1x, so both 5.5 models charge $0.20 to read a cached token. Anthropic says cache reads "make up the majority of agentic and coding work costs", so for an agent with a big stable prefix, the gap between flagship and workhorse is narrower than the headline rates suggest.
A housekeeping note: our Claude Sonnet 5 deep dive, written in June, quotes Sonnet 5 at $3/$15. Anthropic's pricing page today lists $2/$10, so trust the page.
Per-token price is not per-task cost
Per-token price is the number on the rate card. Per-task cost is what you pay for one finished piece of work: a bug fixed, a contract reviewed, an agent run that ends in a pull request. It equals the per-token price multiplied by the tokens the task consumed, and a model can move either factor.
Opus 5.5 moves both. Anthropic says it "costs less per token than Opus 5 and uses fewer tokens per task, which nets out to a 40% drop in costs", and that at default settings it "will cost 40% less than Opus 5 on typical workloads". Note the word typical. Nothing on that page is a claim about your workload.
The 40% is not a price cut. It compounds a rate cut with a change in behaviour. To show the working I need an illustrative workload, and these numbers are mine, not Anthropic's: 10,000 agent tasks a month that between them use 400M cache-read tokens, 50M cache-write tokens (five-minute cache), 50M uncached input tokens and 30M output tokens.
| Line item | Tokens | Opus 5 | Opus 5.5 |
|---|---|---|---|
| Cache reads | 400M | $200.00 | $80.00 |
| Cache writes | 50M | $312.50 | $250.00 |
| Uncached input | 50M | $250.00 | $200.00 |
| Output | 30M | $750.00 | $600.00 |
| Total | $1,512.50 | $1,130.00 |
Step one is the rate card alone. The same tokens cost $1,512.50 on Opus 5 and $1,130.00 on Opus 5.5, a 25.3% cut. That lands between the 20% and the 60%, and where you land depends on your mix. Pure output would see 20%. A cache-read-heavy job would see closer to 60%.
Step two is tokens. Anthropic's 40% needs the tasks themselves to shrink. In this mix, if Opus 5.5 finishes the same tasks on 20% fewer tokens, the bill is 1,130 x 0.8 = $904, which is 40.2% below $1,512.50. With the same token count you get the 25.3% and no more. With more tokens, less. Anthropic has not published its token ratio. The 20% is the figure that makes this mix reproduce the headline, not a number Anthropic gave.
Is there evidence that tokens fall? Anthropic's page is thin on aggregates and rich in anecdotes. In its own tests a 200,000-line codebase audit finished in under three hours, where Opus 5 took over 20 hours and 2.5x as many tokens. Customer figures relayed on the page, which I am paraphrasing, run from Factory's 20 to 25% fewer output tokens to Rogo's roughly 60%. Those are vendor-curated testimonials, not a sample, but a spread of roughly three to one means your number could sit on either side of 40%.
Anthropic also says Opus 5.5 at default effort "beats Opus 5 at max effort for about a fifth of the cost". That compares settings as well as models, so replay your own tasks before believing it.
Why is it cheaper? Anthropic says Opus 5.5 "requires less compute to serve than Opus 5, and its pricing reflects that." It does not say how. Its launch page names no kernel, chip, number format or architecture change. Our working theory, and the published research behind it, is in why frontier AI got cheaper. It is a theory. Nobody at Anthropic has confirmed any of it, and the token savings are a separate question from serving cost.
Sonnet 5.5: same rate card, different bill
Sonnet 5.5 is the mirror image. The rate card is $2 and $10, identical to Sonnet 5. Anthropic says it "requires fewer tokens per task than Sonnet 5", costs "up to 30% less per task", and "batched tool calls together more than Sonnet 5, leading to fewer steps and lower costs". It also says output is 30% or more faster. The saving has to come from tokens and steps, because the rate did not move.
An uncached illustration, again mine. An agent run sends 150,000 input tokens and gets back 15,000. At $2 and $10 that is $0.30 + $0.15 = $0.45. If Sonnet 5.5 does the job on 30% fewer tokens, it costs 0.7 x $0.45 = $0.315. "Up to 30%" is a ceiling, so $0.315 is the best case and $0.45 is the case where nothing improves for you.
One caution if you are coming from Sonnet 4.6 at $3/$15. Sonnet 5.5 looks a third cheaper, but Anthropic's docs say Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text. Price the same text on both and it is about $2.60 against $3, roughly 13% cheaper before any efficiency gain.
The independent view is less flattering. Artificial Analysis (AA), Intelligence Index v4.3.2 at max effort (with fallback), has Sonnet 5.5 at 56 for $7.60 per index task, against Opus 5.5 at 58 for $5.98. Sonnet's price per token is half of Opus's, yet its bill per index task is 27% higher. AA lists 410M output tokens for Sonnet 5.5 at max against an 81M median, and 260M for Opus 5.5, which it calls very verbose. I have not seen AA's token breakdown, so I cannot say exactly why the cost lands there. At high effort the picture reverses: Sonnet 5.5 scores 47 for $1.08, Opus 5.5 scores 54 for $1.82. I have no AA row for Sonnet 5, so none of this tests Anthropic's 30%. It does show that effort can outweigh the rate card.
On Anthropic's own benchmarks (vendor-run), Sonnet 5.5 is a large step up from Sonnet 5 and close to Opus 5.5:
- CursorBench 4.0: 55.5, against 34.1 for Sonnet 5 and 57.8 for Opus 5.5.
- GDPval-AA v2.1 (Elo): 1844, against 1449 and 1846.
- OSWorld 2.1 (partial): 80.1, against 57.0 and 81.8.
On Terminal-Bench 4.0 Sonnet 5.5 scores 70.6 against Opus 5.5's 66.4 (Opus at xhigh). I am not quoting Sonnet 5's figure from that chart, because it looks anomalous and Anthropic has not explained it. It is one win on one benchmark, and it does not make Sonnet 5.5 the better model in general. Anthropic adds that Sonnet 5.5 at low or medium effort beats Sonnet 5's best score "for about a tenth of the cost per task".
Opus 5.5 against Opus 5, with the losses left in
Anthropic ran Opus 5.5 at max effort with adaptive thinking, and Terminal-Bench at xhigh. Not independently reproduced.
| Benchmark (vendor-run) | Opus 5.5 | Opus 5 | Fable 5.1 | GPT-6 Astra |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 66.4 | 52.3 | 55.8 | 57.9 |
| FrontierCode v1.1 | 54.4 | 48.0 | 50.3 | 53.3 |
| CursorBench 4.0 | 57.8 | 46.6 | 51.8 | n.a. |
| GDPval-AA v2.1 (Elo) | 1846 | 1708 | 1735 | 1542 |
| AutomationBench | 40.0 | 26.9 | 31.4 | 41.4 |
| Terminal-Bench-Science 0.1 | 58.7 | 29.0 | 52.6 | 64.6 |
Opus 5.5 beats Opus 5 on every row Anthropic printed, and the biggest gaps are in agentic and terminal work. Terminal-Bench-Science roughly doubles, and AutomationBench rises from 26.9 to 40.0.
Now the losses. GPT-6 Astra is ahead of Opus 5.5 on two rows: AutomationBench (41.4 against 40.0) and Terminal-Bench-Science (64.6 against 58.7). Anthropic prints both. Its footnotes add that production safeguards were on and fallback models completed tasks when they intervened, which "likely lowered some scores".
Effort matters too. On FrontierCode Opus 5.5 scores 54.4 at max and 54.6 at its default medium, so max buys nothing there. On CursorBench it is 57.8 at max and 52.5 at medium. Expect some of your results to look like the second.
Anthropic's cost claims against OpenAI are vendor-run as well. On FrontierCode at default effort, Opus 5.5 "beats GPT-6 Astra at roughly 20% of the cost per task". On CursorBench it beats GPT-5.6 Sol, an older OpenAI model, "by 11 points for about a third of the cost".
What Arena and Artificial Analysis add
Arena's text leaderboard, last updated 25 September with 8,528,723 votes across 409 models, has Claude Opus 5.5 (high) at rank 1 with 1509 +/-12 from 2,307 votes. The next three sit at 1505, 1504 and 1502, so the interval is wider than the lead. Opus 5.5 is on top; it is not statistically separated from #2 to #4. The top six places are all Anthropic models. I could not find Sonnet 5.5 on that board, so I claim no rank for it. Arena's front page also shows an agent board with no update date, where Fable 5.1 (max) has 14.06% and Opus 5.5 (high) 11.84%. Cite arena.ai directly, with the date, because aggregator sites often show stale figures.
Artificial Analysis publishes the Intelligence Index. Every score below is v4.3.2, and the effort setting is named because it moves score and cost together.
- Opus 5.5: 58 at max ($5.98 per index task), 56 at xhigh ($3.46), 54 at high ($1.82), 51 at medium ($1.34).
- Sonnet 5.5: 56 at max ($7.60), 47 at high ($1.08).
- Fable 5.1: 53 at max ($7.63).
Three readings, each my arithmetic on AA's rows. Opus 5.5 at max is 5 points above Fable 5.1 at max for about 22% less per task. Opus 5.5 at high (54) beats Fable 5.1 at max (53) at $1.82 against $7.63. And at default settings, medium for Opus 5.5 and high for Sonnet 5.5, Opus scores 51 against 47 for $1.34 against $1.08: four points for 24% more.
The caveat covering all of it: AA's cost is per index task, on AA's mix. It describes how a model behaves on that mix, not what your traffic costs.
There is also a rival in these numbers. OpenAI's GPT-6.1 Sol, released 29 September at $2/$10, scores 52 at max for $0.72 per task and 50 at high for $0.32 on the same index. That is below Opus 5.5, but 3 points above Sonnet 5.5 at high for under a third of the cost. If Sonnet 5.5 is your candidate, that is the comparison to run on your own tasks. For OpenAI's top tier, see GPT-6 Astra vs Claude Fable 5.1.
Where you can run them, and the dials that move the bill
Anthropic lists both on the Claude API, Amazon Bedrock, Google Cloud's Vertex AI, Microsoft Foundry and Claude Platform on AWS. The API IDs are claude-opus-5-5 and claude-sonnet-5-5, and zero data retention is offered. Press coverage, including VentureBeat, reports both on the Pro, Max, Team and Enterprise plans. I have not seen Anthropic's plan page, so treat that as reported.
Four settings change what you pay.
- Effort. Opus 5.5 uses adaptive thinking that is always on and cannot be disabled, default effort medium. Sonnet 5.5 defaults to high and adds a lowest setting,
between_tools. Default against max is your largest lever. - Batch. Half price on both, for work that can wait.
- Fast mode. A research preview, API only: $8 and $40 per million on Opus 5.5, twice standard, for output Anthropic says is "up to 2.5x" faster. Sonnet 5.5 does not offer it.
- Caching. Reads are $0.20 on both models. A one-hour cache write is $8 on Opus 5.5 and $4 on Sonnet 5.5.
Anthropic said Haiku 5.5 will follow "in the coming weeks". It has not shipped, so I say nothing about its price.
Opus 5.5, Sonnet 5.5 or Fable 5.1
Four rules, by workload.
Hard agentic and long-horizon work: Opus 5.5. Anthropic's table, Arena and AA point the same way. Start at the default medium and move up only if your evals show a gain.
High-volume, moderate-difficulty work: Sonnet 5.5, after a token audit. At $2/$10, or $1/$5 on Batch, it is the cheaper rate for output-heavy work such as summaries, extraction and routine code changes. Measure tokens per task on your own traffic first, and avoid max effort without a reason, given that 410M-token figure. Coming from Sonnet 5 it is a model ID and a regression test, with no price change.
Fable 5.1: keep it only with a reason you can name. It costs 2.5x Opus 5.5 per token. Opus 5.5 is ahead of it on every benchmark Anthropic lists, and ahead on AA at max (58 against 53) for less per task. Fable does lead Arena's agent board, though at a different effort setting. Measured wins on your own evals are a reason to keep it. A sense that it is the premium tier is not.
Simple, high-volume traffic: neither. Opus 5.5's thinking cannot be switched off, and short classification or routing does not need it. Price Haiku 4.5 and non-Anthropic models first.
The method is the same for all four. Pull 30 days of usage split into cache reads, cache writes, uncached input and output, and re-price it on both rate cards. Replay 50 to 100 real tasks at default effort and count the tokens each takes. Then compare cost per finished task, not per million tokens. The token cost calculator handles the rate-card half, and our piece on hidden reasoning-token costs explains why the token half is where estimates go wrong.
Where Swfte sits
The pattern here, a cheaper default model with escalation to a stronger one, is a routing rule. Swfte Connect offers one API to 50+ LLM providers, with smart routing that optimises for cost, latency and availability, and you can bring your own keys or use Swfte's pooled access. We covered the idea in intelligent LLM routing.
Sticker price told you little this week. Opus 5.5 got 20% cheaper and may cost 40% less to run. Sonnet 5.5 did not change price and may cost 30% less. The only number that settles your bill is the one you measure on your own tasks.
Related: Claude Sonnet 5 deep dive, GPT-6 Astra vs Claude Fable 5.1, why frontier AI got cheaper, current rankings on AI model leaderboard.
Sources: Anthropic, Claude Opus 5.5 · Anthropic, Claude Sonnet 5.5 · Claude docs, Opus 5.5 overview · Claude docs, Sonnet 5.5 overview · Claude docs, pricing · Arena, text leaderboard · Artificial Analysis model pages, Intelligence Index v4.3.2, accessed 30 September 2026.
Benchmark figures from Anthropic are vendor-run and not independently reproduced. Worked examples are illustrations with assumed workloads, not measurements. Artificial Analysis and Arena figures were read on 30 September 2026 and will move.