AI Token Cost Calculator

Three common prompt scenarios (short chat, long document summary, agentic tool-use loop) costed across every major AI model using official September 2026 pricing. The bars are dollars per month at the stated request volume.

Token cost calculator

Put your own numbers in.

Monthly cost at a published rate, for the volume and prompt size you actually send.

Estimated monthly cost$55.00Rate used: GPT-4o — $2.5/1M in, $10/1M out

Short customer-support chat

Median support ticket: 800 input tokens (system + ticket body), 250 output tokens (reply). 100K tickets/month.

The spread

Same workload. Gemini 2.0 Flash: $18.00/month. GPT-5.5 Pro: $6,900/month. That is a 383x cost ratio for the same prompt.

GPT-5.5 Pro
$6,900
Claude Opus 4
$3,075
Claude Fable 5
$2,050
Claude Mythos 5
$2,050
GPT-5.5
$1,150
GPT-5.6 Sol
$1,150
Claude Opus 4.7
$1,025
Claude Opus 4.8
$1,025
Claude Opus 5
$1,025
Claude Sonnet 4
$615
Grok 3
$615
Sonar Pro
$615
Kimi K3
$615
Claude Sonnet 4.6
$615
Claude Sonnet 5
$615
Gemini 3.1 Pro
$460
GPT-5.6 Terra
$460
GPT-4o
$450
Command R+
$450
Qwen 3.7 Max
$388
o3
$360
GPT-4.1
$360
Gemini 2.5 Pro
$350
Mistral Large 2
$310
Grok 4.5
$310
Qwen3.8 Max
$310
Qwen 3.6 Plus
$252
GLM-5.2
$222
o3 Mini
$198
Claude 3.5 Haiku
$164
Grok 4.3
$163
GLM-5.1
$155
Gemini 3.6 Flash
$154
Kimi K2.6
$146
Amazon Nova Pro
$144
MiniMax M3
$108
DeepSeek V4 Pro
$102
DeepSeek R1
$98.75
Gemini 3.5 Flash-Lite
$86.50
DeepSeek V3
$49.10
Codestral
$46.50
Qwen 2.5 72B
$46.50
GPT-5.6 Luna
$46.00
Grok 3 Mini
$36.50
Llama 4 Maverick
$31.00
GPT-4o Mini
$27.00
Qwen 2.5 Coder 32B
$23.25
Llama 4 Scout
$22.00
DeepSeek V4 Flash
$18.20
Gemini 2.0 Flash
$18.00

Based on official provider pricing as of 2026-05-06. Excludes prompt-caching, batch, and volume discounts. Self-hosted open-weight models (Gemma 4, Nemotron 3 Nano Omni) excluded from this view because their cost is infrastructure-driven, not token-priced.

Long document summary

40K input tokens (a 30-page PDF), 600 output tokens (executive summary). 5K runs/month.

The spread

Same workload. Gemini 2.0 Flash: $21.20/month. GPT-5.5 Pro: $6,540/month. That is a 308x cost ratio for the same prompt.

GPT-5.5 Pro
$6,540
Claude Opus 4
$3,225
Claude Fable 5
$2,150
Claude Mythos 5
$2,150
GPT-5.5
$1,090
GPT-5.6 Sol
$1,090
Claude Opus 4.7
$1,075
Claude Opus 4.8
$1,075
Claude Opus 5
$1,075
Claude Sonnet 4
$645
Grok 3
$645
Sonar Pro
$645
Kimi K3
$645
Claude Sonnet 4.6
$645
Claude Sonnet 5
$645
GPT-4o
$530
Command R+
$530
Qwen 3.7 Max
$523
Gemini 3.1 Pro
$436
GPT-5.6 Terra
$436
o3
$424
GPT-4.1
$424
Mistral Large 2
$418
Grok 4.5
$418
Qwen3.8 Max
$418
Qwen 3.6 Plus
$297
GLM-5.2
$293
Gemini 2.5 Pro
$280
Grok 4.3
$258
o3 Mini
$233
GLM-5.1
$205
Claude 3.5 Haiku
$172
Amazon Nova Pro
$170
Gemini 3.6 Flash
$161
Kimi K2.6
$156
DeepSeek V4 Pro
$138
MiniMax M3
$127
DeepSeek R1
$117
Gemini 3.5 Flash-Lite
$67.50
Codestral
$62.70
Qwen 2.5 72B
$62.70
Grok 3 Mini
$61.50
DeepSeek V3
$57.30
GPT-5.6 Luna
$43.60
Llama 4 Maverick
$41.80
GPT-4o Mini
$31.80
Qwen 2.5 Coder 32B
$31.35
Llama 4 Scout
$31.20
DeepSeek V4 Flash
$28.84
Gemini 2.0 Flash
$21.20

Based on official provider pricing as of 2026-05-06. Excludes prompt-caching, batch, and volume discounts. Self-hosted open-weight models (Gemma 4, Nemotron 3 Nano Omni) excluded from this view because their cost is infrastructure-driven, not token-priced.

Agentic tool-use loop

6K input tokens per turn (history + tool defs), 400 output tokens, 4 turns per task. 20K tasks/month.

The spread

Same workload. Gemini 2.0 Flash: $60.80/month. GPT-5.5 Pro: $20,160/month. That is a 332x cost ratio for the same prompt.

GPT-5.5 Pro
$20,160
Claude Opus 4
$9,600
Claude Fable 5
$6,400
Claude Mythos 5
$6,400
GPT-5.5
$3,360
GPT-5.6 Sol
$3,360
Claude Opus 4.7
$3,200
Claude Opus 4.8
$3,200
Claude Opus 5
$3,200
Claude Sonnet 4
$1,920
Grok 3
$1,920
Sonar Pro
$1,920
Kimi K3
$1,920
Claude Sonnet 4.6
$1,920
Claude Sonnet 5
$1,920
GPT-4o
$1,520
Command R+
$1,520
Qwen 3.7 Max
$1,440
Gemini 3.1 Pro
$1,344
GPT-5.6 Terra
$1,344
o3
$1,216
GPT-4.1
$1,216
Mistral Large 2
$1,152
Grok 4.5
$1,152
Qwen3.8 Max
$1,152
Gemini 2.5 Pro
$920
Qwen 3.6 Plus
$851
GLM-5.2
$813
Grok 4.3
$680
o3 Mini
$669
GLM-5.1
$569
Claude 3.5 Haiku
$512
Amazon Nova Pro
$486
Gemini 3.6 Flash
$480
Kimi K2.6
$462
DeepSeek V4 Pro
$380
MiniMax M3
$365
DeepSeek R1
$334
Gemini 3.5 Flash-Lite
$224
Codestral
$173
Qwen 2.5 72B
$173
DeepSeek V3
$165
Grok 3 Mini
$160
GPT-5.6 Luna
$134
Llama 4 Maverick
$115
GPT-4o Mini
$91.20
Qwen 2.5 Coder 32B
$86.40
Llama 4 Scout
$84.80
DeepSeek V4 Flash
$76.16
Gemini 2.0 Flash
$60.80

Based on official provider pricing as of 2026-05-06. Excludes prompt-caching, batch, and volume discounts. Self-hosted open-weight models (Gemma 4, Nemotron 3 Nano Omni) excluded from this view because their cost is infrastructure-driven, not token-priced.

How to read these visualizations

The same exact prompt sent to GPT-5.5 Pro versus DeepSeek V4 Flash can differ in cost by more than 200x. That spread is not quality. It is brand, infrastructure, and pricing strategy. Two of the three scenarios above show that even at the "expensive-end" of frontier models, the per-call cost is small; the magic happens at scale, where switching from GPT-5.5 Pro to DeepSeek V4 Pro on a 100K-call/month workload changes the annual line item from millions to tens of thousands.

The right tactic is not "always pick cheap"

For ~80% of production traffic, a mid-tier model (DeepSeek V4 Pro, Gemini 3.1 Pro, Claude Sonnet 4) is the right call. The remaining ~20% (high-stakes reasoning, ambiguous customer queries, agent planning) earns its keep at frontier prices. This is exactly the cascade and Mixture-of-Routers patterns described in our LLM routing deep-dive.

Related calculators

Pricing data sourced from official provider pages and OpenRouter, September 2026. All numbers exclude prompt caching (90% saving on cached input for Anthropic), batch (50% saving on most providers), and committed-use discounts.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.