Updated Aug 4, 2026

Claude API Pricing, August 2026

Definitive Claude API pricing reference. Per-1M-token rates for Opus 5, Sonnet 5, and Haiku 4.5, plus cached-input pricing, batch tier discounts, the Opus 5 tokenizer drift, and a clean comparison vs OpenAI, Gemini, and DeepSeek.

Claude model line, list pricing

ModelInput /1MCached /1MOutput /1MContext
Claude Opus 5

Released 24 Jul 2026 and the current flagship. Drop-in upgrade at Opus 4.8 pricing. Thinking is on by default, and the prompt-cache minimum halves to 512 tokens (from 1024), so shorter prompts now cache. Low and medium effort are unusually strong here — the effective cost is often below the sticker. Fast mode is Claude-API-only at $10/$50. Separate rate-limit pool from the Opus 4.x models.

$5.00$0.50$25.001M
Claude Fable 5

The most capable widely released Anthropic model, for the hardest reasoning and long-horizon agentic work. Always-on thinking (the raw chain of thought is never returned). Requires 30-day data retention, so it is unavailable to zero-retention orgs. Claude Mythos 5 is the same model at the same price, restricted to Project Glasswing.

$10.00$1.00$50.001M
Claude Sonnet 5

The default workhorse tier: ~82% SWE-bench Verified at a third of Opus pricing. An introductory rate of $2/$10 runs through 31 Aug 2026. Note the new tokenizer produces roughly 30% more tokens for the same text than Sonnet 4.6, so re-baseline token budgets rather than reusing old counts.

$3.00$0.30$15.001M
Claude Haiku 4.5

Cheapest Anthropic tier, suited to classification, extraction, and routing. The only current Claude model capped below 1M context (200K) and below 128K output (64K).

$1.00$0.10$5.00200K

All prices in USD per 1M tokens. Cached column is the per-1M rate for cache hits: 90% off list. Cache writes are 25% more expensive than list. Default cache TTL is 5 minutes.

Prompt caching: the largest single lever

Cached input savings, per 1M input tokens (Claude August 2026)

  Opus 5    list $5.00   ████████████████████   cached $0.50  ██   -90%
  Sonnet 5    list $3.00   ████████████           cached $0.30  ██   -90%
  Haiku 4.5   list $0.80   ███                    cached $0.08  ▏    -90%

Production teams that cache the system prompt + tool defs typically
cut Claude bills 40-60% with one config change.

Anthropic's prompt caching is the single largest lever on effective Claude bill in 2026. The mechanism: mark a prefix (system prompt, tool definitions, retrieved context) as cacheable, and any subsequent call within the 5-minute TTL pays 10% of list for those tokens. The break-even on cache writes vs list is around 2 reuses, so any prompt sent more than twice within 5 minutes saves money.

Batch tier, 50% off list

ModelBatch In /1MBatch Out /1MSLA
Claude Opus 5$2.50$12.5024h
Claude Fable 5$5.00$25.0024h
Claude Sonnet 5$1.50$7.5024h
Claude Haiku 4.5$0.50$2.5024h

The Claude Message Batches API gives 50% off list for asynchronous work with a 24-hour turnaround SLA. Stack it with prompt caching for cumulative discounts: a cached batch call on Opus 5 effectively prices at ~$0.25 per 1M cached input, far below most synchronous frontier alternatives.

Tokenizer drift: which upgrades actually cost you more

List price is a poor guide to what an upgrade costs, because Anthropic has changed tokenizers twice and the change is invisible on the rate card. The tokenizer introduced with Opus 4.7 carried through to Opus 4.8, Opus 5, and Fable 5 — so moving between those four is roughly token-neutral. The two migrations that do cost more are coming from Opus 4.6 or earlier, and moving from Sonnet 4.6 to Sonnet 5.

Effective cost change by migration path (identical input text)

  Migration                    List price      Effective change
  Opus 4.8 / 4.7  -> Opus 5    $5/$25 per 1M   ~0% (same tokenizer)
  Opus 4.6        -> Opus 5    $5/$25 per 1M   +0% to +35% (varies)
  Sonnet 4.6      -> Sonnet 5  $3/$15 per 1M   ~+30% (new tokenizer)
  Sonnet 5 intro rate to 31 Aug 2026: $2/$10 per 1M

One more adder that is easy to miss on Opus 5: thinking is on by default, where on Opus 4.8 and 4.7 omitting the parameter meant no thinking at all. Thinking tokens bill as output, so a workload that was implicitly thinking-free can get materially more expensive on identical code. Set effort to low or medium — both are unusually strong on this model — or disable thinking explicitly at effort high or below.

See our per-million-tokens true cost writeup for the full hidden-adders breakdown.

Claude vs OpenAI vs Gemini vs DeepSeek

VendorModelIn /1MOut /1MCacheBatch
AnthropicClaude Opus 5$5.00$25.00YesYes
OpenAIGPT-5.6 Sol$5.00$30.00YesYes
OpenAIGPT-5.6 Terra$2.00$12.00YesYes
OpenAIGPT-5.6 Luna$0.20$1.20YesYes
GoogleGemini 3.1 Pro$2.00$12.00YesYes
GoogleGemini 3.6 Flash$1.50$7.50YesYes
xAIGrok 4.5$2.00$6.00Yes
MoonshotKimi K3 (open weight)$3.00$15.00Yes
Z.aiGLM-5.2 (MIT)$1.40$4.40Yes
DeepSeekDeepSeek V4 Pro (MIT)$0.43$0.87Yes
DeepSeekDeepSeek V4 Flash (MIT)$0.14$0.28Yes
Output cost per 1M tokens (frontier tier, August 2026)

  GPT-5.5 Pro       $180.00  ████████████████████████████  most expensive
  GPT-5.5            $30.00  █████                         standard tier
  Claude Opus 5    $25.00  ████                          coding leader
  Sonnet 5           $15.00  ██                            workhorse
  Gemini 3.1 Pro     $10.50  █▌                            text Arena #1
  DeepSeek V4 Pro     $3.48  ▌                             open weights
  DeepSeek V4 Flash   $0.28  ▏                             cheapest tier

Spread between cheapest and most expensive: ~643x

What to do this quarter

  1. Turn on prompt caching everywhere. System prompts and tool definitions are the same on every call: they are pure cache wins. Expect 40-60% bill reduction with one config change.
  2. Move asynchronous workloads to batch. Anything that can wait 24 hours should be on the batch tier at 50% off. Stacks with caching.
  3. Re-baseline tokens on every Anthropic launch. The 4.6 to 4.7 tokenizer change added 35% to effective cost with no list-price change. Treat tokenizer churn as a price-change event.
  4. Use Sonnet 5 as the default, Opus 5 by exception. Sonnet 5 covers 70% of typical agentic and chat workloads at one-fifth the price. Reserve Opus 5 for hardest-task coding and reasoning where it earns its premium.
  5. Pair with a cheap-tier fallback. Cascade trivial traffic to DeepSeek V4 Flash or Gemini 3.1 Pro. Quality delta on simple work is rounding error; cost delta is 10-200x.
  6. Negotiate volume committed-spend discounts. Anthropic offers committed-use pricing above ~$50K/mo. The public list is the floor.
  7. Track the bill, not the list. Effective rate = bill divided by external prompts handled. That is the only number that matters at quarter end.

Related reading

Teams running Claude alongside other providers typically front the API with Swfte Connect to keep the OpenAI-compatible surface, prompt caching, and per-route fallback in one place rather than re-implementing it per vendor.

Sources: official Anthropic pricing page, May 2026-08-04. Tokenizer drift figures from internal Swfte Connect telemetry on representative SaaS workload mixes.