Claude API Pricing, August 2026
Definitive Claude API pricing reference. Per-1M-token rates for Opus 5, Sonnet 5, and Haiku 4.5, plus cached-input pricing, batch tier discounts, the Opus 5 tokenizer drift, and a clean comparison vs OpenAI, Gemini, and DeepSeek.
Claude model line, list pricing
| Model | Input /1M | Cached /1M | Output /1M | Context |
|---|---|---|---|---|
| Claude Opus 5 Released 24 Jul 2026 and the current flagship. Drop-in upgrade at Opus 4.8 pricing. Thinking is on by default, and the prompt-cache minimum halves to 512 tokens (from 1024), so shorter prompts now cache. Low and medium effort are unusually strong here — the effective cost is often below the sticker. Fast mode is Claude-API-only at $10/$50. Separate rate-limit pool from the Opus 4.x models. | $5.00 | $0.50 | $25.00 | 1M |
| Claude Fable 5 The most capable widely released Anthropic model, for the hardest reasoning and long-horizon agentic work. Always-on thinking (the raw chain of thought is never returned). Requires 30-day data retention, so it is unavailable to zero-retention orgs. Claude Mythos 5 is the same model at the same price, restricted to Project Glasswing. | $10.00 | $1.00 | $50.00 | 1M |
| Claude Sonnet 5 The default workhorse tier: ~82% SWE-bench Verified at a third of Opus pricing. An introductory rate of $2/$10 runs through 31 Aug 2026. Note the new tokenizer produces roughly 30% more tokens for the same text than Sonnet 4.6, so re-baseline token budgets rather than reusing old counts. | $3.00 | $0.30 | $15.00 | 1M |
| Claude Haiku 4.5 Cheapest Anthropic tier, suited to classification, extraction, and routing. The only current Claude model capped below 1M context (200K) and below 128K output (64K). | $1.00 | $0.10 | $5.00 | 200K |
All prices in USD per 1M tokens. Cached column is the per-1M rate for cache hits: 90% off list. Cache writes are 25% more expensive than list. Default cache TTL is 5 minutes.
Prompt caching: the largest single lever
Cached input savings, per 1M input tokens (Claude August 2026) Opus 5 list $5.00 ████████████████████ cached $0.50 ██ -90% Sonnet 5 list $3.00 ████████████ cached $0.30 ██ -90% Haiku 4.5 list $0.80 ███ cached $0.08 ▏ -90% Production teams that cache the system prompt + tool defs typically cut Claude bills 40-60% with one config change.
Anthropic's prompt caching is the single largest lever on effective Claude bill in 2026. The mechanism: mark a prefix (system prompt, tool definitions, retrieved context) as cacheable, and any subsequent call within the 5-minute TTL pays 10% of list for those tokens. The break-even on cache writes vs list is around 2 reuses, so any prompt sent more than twice within 5 minutes saves money.
Batch tier, 50% off list
| Model | Batch In /1M | Batch Out /1M | SLA |
|---|---|---|---|
| Claude Opus 5 | $2.50 | $12.50 | 24h |
| Claude Fable 5 | $5.00 | $25.00 | 24h |
| Claude Sonnet 5 | $1.50 | $7.50 | 24h |
| Claude Haiku 4.5 | $0.50 | $2.50 | 24h |
The Claude Message Batches API gives 50% off list for asynchronous work with a 24-hour turnaround SLA. Stack it with prompt caching for cumulative discounts: a cached batch call on Opus 5 effectively prices at ~$0.25 per 1M cached input, far below most synchronous frontier alternatives.
Tokenizer drift: which upgrades actually cost you more
List price is a poor guide to what an upgrade costs, because Anthropic has changed tokenizers twice and the change is invisible on the rate card. The tokenizer introduced with Opus 4.7 carried through to Opus 4.8, Opus 5, and Fable 5 — so moving between those four is roughly token-neutral. The two migrations that do cost more are coming from Opus 4.6 or earlier, and moving from Sonnet 4.6 to Sonnet 5.
Effective cost change by migration path (identical input text) Migration List price Effective change Opus 4.8 / 4.7 -> Opus 5 $5/$25 per 1M ~0% (same tokenizer) Opus 4.6 -> Opus 5 $5/$25 per 1M +0% to +35% (varies) Sonnet 4.6 -> Sonnet 5 $3/$15 per 1M ~+30% (new tokenizer) Sonnet 5 intro rate to 31 Aug 2026: $2/$10 per 1M
One more adder that is easy to miss on Opus 5: thinking is on by default, where on Opus 4.8 and 4.7 omitting the parameter meant no thinking at all. Thinking tokens bill as output, so a workload that was implicitly thinking-free can get materially more expensive on identical code. Set effort to low or medium — both are unusually strong on this model — or disable thinking explicitly at effort high or below.
See our per-million-tokens true cost writeup for the full hidden-adders breakdown.
Claude vs OpenAI vs Gemini vs DeepSeek
| Vendor | Model | In /1M | Out /1M | Cache | Batch |
|---|---|---|---|---|---|
| Anthropic | Claude Opus 5 | $5.00 | $25.00 | Yes | Yes |
| OpenAI | GPT-5.6 Sol | $5.00 | $30.00 | Yes | Yes |
| OpenAI | GPT-5.6 Terra | $2.00 | $12.00 | Yes | Yes |
| OpenAI | GPT-5.6 Luna | $0.20 | $1.20 | Yes | Yes |
| Gemini 3.1 Pro | $2.00 | $12.00 | Yes | Yes | |
| Gemini 3.6 Flash | $1.50 | $7.50 | Yes | Yes | |
| xAI | Grok 4.5 | $2.00 | $6.00 | Yes | — |
| Moonshot | Kimi K3 (open weight) | $3.00 | $15.00 | Yes | — |
| Z.ai | GLM-5.2 (MIT) | $1.40 | $4.40 | Yes | — |
| DeepSeek | DeepSeek V4 Pro (MIT) | $0.43 | $0.87 | Yes | — |
| DeepSeek | DeepSeek V4 Flash (MIT) | $0.14 | $0.28 | Yes | — |
Output cost per 1M tokens (frontier tier, August 2026) GPT-5.5 Pro $180.00 ████████████████████████████ most expensive GPT-5.5 $30.00 █████ standard tier Claude Opus 5 $25.00 ████ coding leader Sonnet 5 $15.00 ██ workhorse Gemini 3.1 Pro $10.50 █▌ text Arena #1 DeepSeek V4 Pro $3.48 ▌ open weights DeepSeek V4 Flash $0.28 ▏ cheapest tier Spread between cheapest and most expensive: ~643x
What to do this quarter
- Turn on prompt caching everywhere. System prompts and tool definitions are the same on every call: they are pure cache wins. Expect 40-60% bill reduction with one config change.
- Move asynchronous workloads to batch. Anything that can wait 24 hours should be on the batch tier at 50% off. Stacks with caching.
- Re-baseline tokens on every Anthropic launch. The 4.6 to 4.7 tokenizer change added 35% to effective cost with no list-price change. Treat tokenizer churn as a price-change event.
- Use Sonnet 5 as the default, Opus 5 by exception. Sonnet 5 covers 70% of typical agentic and chat workloads at one-fifth the price. Reserve Opus 5 for hardest-task coding and reasoning where it earns its premium.
- Pair with a cheap-tier fallback. Cascade trivial traffic to DeepSeek V4 Flash or Gemini 3.1 Pro. Quality delta on simple work is rounding error; cost delta is 10-200x.
- Negotiate volume committed-spend discounts. Anthropic offers committed-use pricing above ~$50K/mo. The public list is the floor.
- Track the bill, not the list. Effective rate = bill divided by external prompts handled. That is the only number that matters at quarter end.
Related reading
- AI Model Leaderboard: quality vs price across all providers
- Per Million Tokens True Cost : the hidden adders pushing your bill 1.5-3x above list
- Token Cost Calculator : interactive cost estimator
- OpenRouter vs Swfte Pricing : enterprise procurement comparison
Teams running Claude alongside other providers typically front the API with Swfte Connect to keep the OpenAI-compatible surface, prompt caching, and per-route fallback in one place rather than re-implementing it per vendor.
Sources: official Anthropic pricing page, May 2026-08-04. Tokenizer drift figures from internal Swfte Connect telemetry on representative SaaS workload mixes.