Understand the rate.
Then model your workload.
Compare input and output token prices across every major provider. These are model rates, separate from Swfte product subscriptions and platform fees.
56
Models Tracked
17
Providers
$0.10
Cheapest Input
300x
Price Range
56 of 56 models
Model rates as recorded on Sep 15, 2026. Review dates and assumptions before making a decision. Swfte product subscriptions and Connect platform fees are priced separately.
| Explore | ||||||||
|---|---|---|---|---|---|---|---|---|
| Gemma 4 27B OSS | Self-host | Self-host | Self-host | 75 | 0.0 | 128K | Profile | |
| Nemotron 3 Nano Omni OSS | NVIDIA | Self-host | Self-host | Self-host | 76 | 0.0 | 256K | Profile |
| Qwen3.8 27B OSS | Alibaba Cloud | Self-host | Self-host | Self-host | 80 | 0.0 | 262K | Profile |
| Nemotron 3.5 Lightning OSS | NVIDIA | Self-host | Self-host | Self-host | 79 | 0.0 | 1M | Profile |
| Hunyuan Hy3 OSS | Tencent | Self-host | Self-host | Self-host | 76 | 0.0 | 256K | Profile |
| Ling-3.0-Flash OSS | Ant Group | Self-host | Self-host | Self-host | 74 | 0.0 | 256K | Profile |
| DeepSeek V4 Flash OSS | DeepSeek | $0.14 | $0.28 | $0.21 | 84 | 400.0 | 1M | Profile |
| Gemini 2.0 Flash | $0.10 | $0.40 | $0.25 | 74 | 296.0 | 1M | Profile | |
| Llama 4 Scout OSS | Meta | $0.15 | $0.40 | $0.28 | 71 | 258.2 | 10M | Profile |
| Qwen 2.5 Coder 32B OSS | Alibaba Cloud | $0.15 | $0.45 | $0.30 | 74 | 246.7 | 131K | Profile |
| GPT-4o Mini | OpenAI | $0.15 | $0.60 | $0.38 | 72 | 192.0 | 128K | Profile |
| Llama 4 Maverick OSS | Meta | $0.20 | $0.60 | $0.40 | 80 | 200.0 | 1M | Profile |
| Grok 3 Mini | xAI | $0.30 | $0.50 | $0.40 | 78 | 195.0 | 131K | Profile |
| Codestral | Mistral AI | $0.30 | $0.90 | $0.60 | 76 | 126.7 | 256K | Profile |
| Qwen 2.5 72B OSS | Alibaba Cloud | $0.30 | $0.90 | $0.60 | 80 | 133.3 | 131K | Profile |
| DeepSeek V3 OSS | DeepSeek | $0.27 | $1.10 | $0.69 | 86 | 125.5 | 128K | Profile |
| GPT-5.6 Luna −80% | OpenAI | $0.20 | $1.20 | $0.70 | 83 | 118.6 | 1M | Profile |
| DeepSeek V4 Pro OSS | DeepSeek | $0.66 | $1.98 | $1.32 | 89 | 67.4 | 1M | Profile |
| DeepSeek R1 OSS | DeepSeek | $0.55 | $2.19 | $1.37 | 91 | 66.4 | 128K | Profile |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | $1.40 | 79 | 56.4 | 1M | Profile | |
| MiniMax M3 OSS | MiniMax | $0.60 | $2.40 | $1.50 | 89 | 59.3 | 1M | Profile |
| Grok 4.3 | xAI | $1.25 | $2.50 | $1.88 | 93 | 49.6 | 1M | Profile |
| Amazon Nova Pro | Amazon | $0.80 | $3.20 | $2.00 | 70 | 35.0 | 300K | Profile |
| GLM-5.1 OSS | Z.ai (Zhipu AI) | $0.98 | $3.08 | $2.03 | 88 | 43.3 | 200K | Profile |
| Kimi K2.6 | Moonshot AI | $0.73 | $3.49 | $2.11 | 92 | 43.6 | 256K | Profile |
| Gemini 3.6 Flash | $0.75 | $3.75 | $2.25 | 90 | 40.0 | 1M | Profile | |
| Claude 3.5 Haiku −20% | Anthropic | $0.80 | $4.00 | $2.40 | 75 | 31.3 | 200K | Profile |
| o3 Mini | OpenAI | $1.10 | $4.40 | $2.75 | 88 | 32.0 | 200K | Profile |
| GLM-5.2 OSS | Z.ai (Zhipu AI) | $1.40 | $4.40 | $2.90 | 91 | 31.4 | 1M | Profile |
| Qwen 3.6 Plus | Alibaba Cloud | $1.40 | $5.60 | $3.50 | 86 | 24.6 | 256K | Profile |
| Mistral Large 2 −33% | Mistral AI | $2.00 | $6.00 | $4.00 | 79 | 19.8 | 128K | Profile |
| Grok 4.5 | xAI | $2.00 | $6.00 | $4.00 | 94 | 23.5 | 500K | Profile |
| Qwen3.8 Max OSS | Alibaba Cloud | $2.00 | $6.00 | $4.00 | 96 | 24.0 | 1M | Profile |
| o3 −80% | OpenAI | $2.00 | $8.00 | $5.00 | 94 | 18.8 | 200K | Profile |
| GPT-4.1 | OpenAI | $2.00 | $8.00 | $5.00 | 89 | 17.8 | 1M | Profile |
| Qwen 3.7 Max | Alibaba Cloud | $2.50 | $7.50 | $5.00 | 94 | 18.8 | 1M | Profile |
| Gemini 2.5 Pro | $1.25 | $10.00 | $5.63 | 92 | 16.4 | 1M | Profile | |
| GPT-4o | OpenAI | $2.50 | $10.00 | $6.25 | 85 | 13.6 | 128K | Profile |
| Command R+ | Cohere | $2.50 | $10.00 | $6.25 | 68 | 10.9 | 128K | Profile |
| Gemini 3.1 Pro | $2.00 | $12.00 | $7.00 | 96 | 13.7 | 1M | Profile | |
| GPT-5.6 Terra −20% | OpenAI | $2.00 | $12.00 | $7.00 | 93 | 13.3 | 1M | Profile |
| Claude Sonnet 4 | Anthropic | $3.00 | $15.00 | $9.00 | 88 | 9.8 | 200K | Profile |
| Grok 3 | xAI | $3.00 | $15.00 | $9.00 | 87 | 9.7 | 131K | Profile |
| Sonar Pro | Perplexity | $3.00 | $15.00 | $9.00 | 78 | 8.7 | 200K | Profile |
| Kimi K3 OSS | Moonshot AI | $3.00 | $15.00 | $9.00 | 97 | 10.8 | 1M | Profile |
| Claude Sonnet 4.6 | Anthropic | $3.00 | $15.00 | $9.00 | 90 | 10.0 | 1M | Profile |
| Claude Sonnet 5 | Anthropic | $3.00 | $15.00 | $9.00 | 93 | 10.3 | 1M | Profile |
| Claude Opus 4.7 | Anthropic | $5.00 | $25.00 | $15.00 | 96 | 6.4 | 1M | Profile |
| Claude Opus 4.8 | Anthropic | $5.00 | $25.00 | $15.00 | 98 | 6.5 | 1M | Profile |
| Claude Opus 5 | Anthropic | $5.00 | $25.00 | $15.00 | 99 | 6.6 | 1M | Profile |
| GPT-5.5 | OpenAI | $5.00 | $30.00 | $17.50 | 97 | 5.5 | 1M | Profile |
| GPT-5.6 Sol | OpenAI | $5.00 | $30.00 | $17.50 | 98 | 5.6 | 1M | Profile |
| Claude Fable 5 | Anthropic | $10.00 | $50.00 | $30.00 | 100 | 3.3 | 1M | Profile |
| Claude Mythos 5 | Anthropic | $10.00 | $50.00 | $30.00 | 100 | 3.3 | 1M | Profile |
| Claude Opus 4 | Anthropic | $15.00 | $75.00 | $45.00 | 91 | 2.0 | 200K | Profile |
| GPT-5.5 Pro | OpenAI | $30.00 | $180.00 | $105.00 | 96 | 0.9 | 1M | Profile |
Blended is the average of the input and output rate per 1M tokens. Quality is a composite benchmark score out of 100. Value is quality per dollar — higher is better. Self-hosted models carry no per-token vendor rate.
Estimate Your Monthly Cost
Monthly cost estimate
Enter your typical request shape. Costs below are projected over one month, based on current public list-price API rates.
Cheapest
DeepSeek V4 Flash
$15.40
per month at this volume
Best value (quality ≥ 80)
DeepSeek V4 Flash · Q 84
$15.40
per month at this volume
Most expensive
GPT-5.5 Pro
$6900.00
per month at this volume
Save 30-60% with Mixture-of-Routers
Most production traffic is mixed-difficulty. Send the easy 60% to a cheap model and the hard 10% to a frontier model: same quality, fraction of the cost.
Full breakdown by model
Sorted cheapest to most expensive
| Model | Cost / request | Input cost / mo | Output cost / mo | Total / mo |
|---|---|---|---|---|
DeepSeek V4 Flash $0.14 in / $0.28 out per 1M | $0.000154 | $7.00 | $8.40 | $15.40 |
Gemini 2.0 Flash $0.1 in / $0.4 out per 1M | $0.000170 | $5.00 | $12.00 | $17.00 |
Llama 4 Scout $0.15 in / $0.4 out per 1M | $0.000195 | $7.50 | $12.00 | $19.50 |
Qwen 2.5 Coder 32B $0.15 in / $0.45 out per 1M | $0.000210 | $7.50 | $13.50 | $21.00 |
GPT-4o Mini $0.15 in / $0.6 out per 1M | $0.000255 | $7.50 | $18.00 | $25.50 |
Llama 4 Maverick $0.2 in / $0.6 out per 1M | $0.000280 | $10.00 | $18.00 | $28.00 |
Grok 3 Mini $0.3 in / $0.5 out per 1M | $0.000300 | $15.00 | $15.00 | $30.00 |
Codestral $0.3 in / $0.9 out per 1M | $0.000420 | $15.00 | $27.00 | $42.00 |
Qwen 2.5 72B $0.3 in / $0.9 out per 1M | $0.000420 | $15.00 | $27.00 | $42.00 |
GPT-5.6 Luna $0.2 in / $1.2 out per 1M | $0.000460 | $10.00 | $36.00 | $46.00 |
DeepSeek V3 $0.27 in / $1.1 out per 1M | $0.000465 | $13.50 | $33.00 | $46.50 |
Gemini 3.5 Flash-Lite $0.3 in / $2.5 out per 1M | $0.000900 | $15.00 | $75.00 | $90.00 |
DeepSeek V4 Pro $0.66 in / $1.98 out per 1M | $0.000924 | $33.00 | $59.40 | $92.40 |
DeepSeek R1 $0.55 in / $2.19 out per 1M | $0.000932 | $27.50 | $65.70 | $93.20 |
MiniMax M3 $0.6 in / $2.4 out per 1M | $0.001020 | $30.00 | $72.00 | $102.00 |
Amazon Nova Pro $0.8 in / $3.2 out per 1M | $0.001360 | $40.00 | $96.00 | $136.00 |
Grok 4.3 $1.25 in / $2.5 out per 1M | $0.001375 | $62.50 | $75.00 | $137.50 |
Kimi K2.6 $0.73 in / $3.49 out per 1M | $0.001412 | $36.50 | $104.70 | $141.20 |
GLM-5.1 $0.98 in / $3.08 out per 1M | $0.001414 | $49.00 | $92.40 | $141.40 |
Gemini 3.6 Flash $0.75 in / $3.75 out per 1M | $0.001500 | $37.50 | $112.50 | $150.00 |
Claude 3.5 Haiku $0.8 in / $4 out per 1M | $0.001600 | $40.00 | $120.00 | $160.00 |
o3 Mini $1.1 in / $4.4 out per 1M | $0.001870 | $55.00 | $132.00 | $187.00 |
GLM-5.2 $1.4 in / $4.4 out per 1M | $0.002020 | $70.00 | $132.00 | $202.00 |
Qwen 3.6 Plus $1.4 in / $5.6 out per 1M | $0.002380 | $70.00 | $168.00 | $238.00 |
Mistral Large 2 $2 in / $6 out per 1M | $0.002800 | $100.00 | $180.00 | $280.00 |
Grok 4.5 $2 in / $6 out per 1M | $0.002800 | $100.00 | $180.00 | $280.00 |
Qwen3.8 Max $2 in / $6 out per 1M | $0.002800 | $100.00 | $180.00 | $280.00 |
o3 $2 in / $8 out per 1M | $0.003400 | $100.00 | $240.00 | $340.00 |
GPT-4.1 $2 in / $8 out per 1M | $0.003400 | $100.00 | $240.00 | $340.00 |
Qwen 3.7 Max $2.5 in / $7.5 out per 1M | $0.003500 | $125.00 | $225.00 | $350.00 |
Gemini 2.5 Pro $1.25 in / $10 out per 1M | $0.003625 | $62.50 | $300.00 | $362.50 |
GPT-4o $2.5 in / $10 out per 1M | $0.004250 | $125.00 | $300.00 | $425.00 |
Command R+ $2.5 in / $10 out per 1M | $0.004250 | $125.00 | $300.00 | $425.00 |
Gemini 3.1 Pro $2 in / $12 out per 1M | $0.004600 | $100.00 | $360.00 | $460.00 |
GPT-5.6 Terra $2 in / $12 out per 1M | $0.004600 | $100.00 | $360.00 | $460.00 |
Claude Sonnet 4 $3 in / $15 out per 1M | $0.006000 | $150.00 | $450.00 | $600.00 |
Grok 3 $3 in / $15 out per 1M | $0.006000 | $150.00 | $450.00 | $600.00 |
Sonar Pro $3 in / $15 out per 1M | $0.006000 | $150.00 | $450.00 | $600.00 |
Kimi K3 $3 in / $15 out per 1M | $0.006000 | $150.00 | $450.00 | $600.00 |
Claude Sonnet 4.6 $3 in / $15 out per 1M | $0.006000 | $150.00 | $450.00 | $600.00 |
Claude Sonnet 5 $3 in / $15 out per 1M | $0.006000 | $150.00 | $450.00 | $600.00 |
Claude Opus 4.7 $5 in / $25 out per 1M | $0.0100 | $250.00 | $750.00 | $1000.00 |
Claude Opus 4.8 $5 in / $25 out per 1M | $0.0100 | $250.00 | $750.00 | $1000.00 |
Claude Opus 5 $5 in / $25 out per 1M | $0.0100 | $250.00 | $750.00 | $1000.00 |
GPT-5.5 $5 in / $30 out per 1M | $0.0115 | $250.00 | $900.00 | $1150.00 |
GPT-5.6 Sol $5 in / $30 out per 1M | $0.0115 | $250.00 | $900.00 | $1150.00 |
Claude Fable 5 $10 in / $50 out per 1M | $0.0200 | $500.00 | $1500.00 | $2000.00 |
Claude Mythos 5 $10 in / $50 out per 1M | $0.0200 | $500.00 | $1500.00 | $2000.00 |
Claude Opus 4 $15 in / $75 out per 1M | $0.0300 | $750.00 | $2250.00 | $3000.00 |
GPT-5.5 Pro $30 in / $180 out per 1M | $0.0690 | $1500.00 | $5400.00 | $6900.00 |
Gemma 4 27B Self-hostOpen weights (Apache 2.0) , token cost is $0; infra cost depends on hardware | — | — | — | Self-host |
Nemotron 3 Nano Omni Self-hostOpen weights (NVIDIA Open Model License) , token cost is $0; infra cost depends on hardware | — | — | — | Self-host |
Qwen3.8 27B Self-hostOpen weights (Apache 2.0) , token cost is $0; infra cost depends on hardware | — | — | — | Self-host |
Nemotron 3.5 Lightning Self-hostOpen weights (OpenMDW-1.1) , token cost is $0; infra cost depends on hardware | — | — | — | Self-host |
Hunyuan Hy3 Self-hostOpen weights (Apache 2.0) , token cost is $0; infra cost depends on hardware | — | — | — | Self-host |
Ling-3.0-Flash Self-hostOpen weights (MIT) , token cost is $0; infra cost depends on hardware | — | — | — | Self-host |
List-price estimate. Real bills typically run 1.3-1.7x higher after retries, system-prompt re-sends, and tool-call round-trips. See per-million-tokens true cost for the adders.
Recent Price Changes
DeepSeek V4 Pro
Aug 13, 2026
$0.66 / $1.98
Qwen3.8 Max
Aug 3, 2026
$2 / $6
GPT-5.6 Terra
Jul 30, 2026
$2 / $12
GPT-5.6 Luna
Jul 30, 2026
$0.2 / $1.2
Gemini 3.6 Flash
Jul 21, 2026
$0.75 / $3.75
MiniMax M3
Jun 8, 2026
$0.6 / $2.4
DeepSeek V4 Pro
May 31, 2026
$0.435 / $0.87
Kimi K2.6
May 25, 2026
$0.73 / $3.49
GLM-5.1
May 25, 2026
$0.98 / $3.08
o3
Jun 10, 2025
$2 / $8
Understanding AI API Pricing in 2026
AI model pricing has undergone a dramatic transformation. Since GPT-4 launched in March 2023 at $30 per million input tokens, prices have fallen by over 90%: driven by competition from Anthropic, Google, and open-source challengers like DeepSeek and Meta's Llama.
Today's pricing market spans a 150x range: from Google's Gemini 2.0 Flash at $0.10/1M input tokens to Claude Opus 4 at $15/1M tokens. The key insight is that price doesn't always correlate with quality: DeepSeek V3 delivers 86% quality at just $0.27/1M tokens, while some premium models charge 50x more for marginal quality gains.
How to Optimize AI API Costs
The most effective strategy is model routing: sending simple queries to cheap, fast models and complex queries to premium models. A gateway like Swfte Connect automates this, typically reducing costs by 30-60% without sacrificing quality.
Other strategies include: using cached input pricing (offered by Google and DeepSeek), batching requests to reduce per-call overhead, and using open-source models for predictable workloads where you can self-host.
Pricing Trends to Watch
- Price compression is accelerating: OpenAI cut GPT-5.6 Luna 80% and Terra 20% on 30 July 2026, citing inference work that reduced end-to-end serving cost by 20% and improved token-generation efficiency by more than 15%
- DeepSeek sets the floor: V4 Pro launched at $1.74/$3.48, was discounted 75%, and had that cut made permanent on 31 May 2026 — so frontier-class quality now costs $0.435/$0.87. Every other lab is now pricing against that number
- Open-weight pressure, not open-source pressure: the squeeze is coming from Chinese labs — DeepSeek, Moonshot, Z.ai, Alibaba, MiniMax — not Meta, which cancelled its open-frontier plans and shipped the closed Muse Spark instead
- Reasoning premium: Models with extended thinking cost more per request because reasoning tokens bill as output. On Kimi K3 thinking cannot be disabled at any setting
- Cached pricing is the real discount: cache reads run 10% of input at OpenAI and Google, and DeepSeek cut its cache-hit rate to a tenth of launch price on 26 April 2026 — $0.0028 per 1M tokens on V4 Flash, effectively free