AI Model Directory
Detailed specs, pricing, and benchmarks for every major AI model.
Claude Fable 5
Anthropic's June 9 2026 flagship and the first generally available model in the Mythos class — the new #1 overall. State-of-the-art on nearly every tested benchmark: SWE-bench Pro 80.3% (vs Opus 4.8 69.2%, GPT-5.5 58.6%), FrontierCode Diamond 29.3%, OSWorld-Verified computer use 85.0%, and a leading GDPval-AA knowledge-work Elo of 1932. Routes high-risk cyber/bio prompts to Opus 4.8 as a safety fallback. $10/$50 pricing (2x Opus 4.8) with a 90% prompt-caching discount; exact context window and max output were not officially disclosed at launch.
100
Quality
$30.00
Blended/1M
58
tok/s
Claude Opus 4.8
Anthropic's May 28 2026 flagship and the new #1 on the Artificial Analysis Intelligence Index at 61.4 (+4.1 over Opus 4.7, +1.2 ahead of GPT-5.5). SWE-bench Verified 88.6%, SWE-bench Pro 69.2%, Terminal-Bench 2.1 74.6%, and a leading 49.8% on Humanity's Last Exam. Strongest computer-use/browser-agent model tested (Online-Mind2Web 84%). Same $5/$25 pricing as Opus 4.7; fast mode is ~2.5x faster and ~3x cheaper.
99
Quality
$15.00
Blended/1M
72
tok/s
GPT-5.5 Pro
High-compute variant of GPT-5.5 with extended thinking. 6x the price for mid-teens AAII uplift on hard reasoning workloads.
98
Quality
$105.00
Blended/1M
68
tok/s
GPT-5.6
OpenAI's July 8 2026 flagship — a point release over GPT-5.5 that closes most of the gap to Claude Opus 4.8. Artificial Analysis Intelligence Index 61.0 (up from GPT-5.5's 60.2), SWE-bench Verified 88.1%, SWE-bench Pro 66.4%, Terminal-Bench 2.1 72.9%. The headline change is agentic reliability: OpenAI reports far fewer dropped tool calls on long multi-step runs. Pricing is unchanged from GPT-5.5 at $5/$30 with a 90% cached-input discount, so the quality gain is effectively free for existing users.
98
Quality
$17.50
Blended/1M
96
tok/s
GPT-5.5
OpenAI's "Spud" — the first fully retrained base model since GPT-4.5. 1M context window; tops the Artificial Analysis Intelligence Index at 59-60 and the LMArena text leaderboard at ~1506 Elo. Shipped April 23 2026, on AWS Bedrock April 28.
97
Quality
$17.50
Blended/1M
70
tok/s
Gemini 3.2 Pro
Google's July 2 2026 flagship and a clear step over Gemini 3.1 Pro. Artificial Analysis Intelligence Index 59.6, SWE-bench Verified 86.4%, and a class-leading 2M-token context that holds accuracy deeper into the window than any rival on Google's own long-context evals. Still the value pick among frontier proprietary models at $2/$12 (unchanged from 3.1 Pro), with native audio and video understanding. Best all-round choice when the workload is long documents, multimodal input, or high request volume where price matters.
97
Quality
$7.00
Blended/1M
128
tok/s
Claude Opus 4.7
Anthropic's April 16 2026 flagship. SWE-bench Verified 87.6%, SWE-bench Pro 64.3%, GPQA Diamond 94.2%. Sits in the LMArena top tier at ~1503 (1505 thinking). Same $5/$25 pricing as Opus 4.6.
96
Quality
$15.00
Blended/1M
68
tok/s
Gemini 3.1 Pro
Tied at the top of the LMArena text leaderboard May 2026 @ ~1505 Elo. GPQA Diamond 94.3%, MMLU-Pro 91.0% — current scientific reasoning leader. 1M token context; fastest frontier model at ~131 tok/s.
96
Quality
$7.00
Blended/1M
131
tok/s
o3
OpenAI's most powerful reasoning model. State-of-the-art on MATH, coding, and science benchmarks.
94
Quality
$25.00
Blended/1M
68
tok/s
Qwen 3.7 Max
Alibaba's May 20 2026 proprietary flagship, unveiled at the Alibaba Cloud Summit in Hangzhou. Highest-ranked Chinese model on the Artificial Analysis Intelligence Index at 56.6 (#5 overall, +4.8 over Qwen 3.6 Max Preview). 1M context; SWE-bench Pro 60.6, Terminal-Bench 2.0 69.7, GPQA Diamond 92.4, and a table-leading 97.1 on HMMT Feb 2026 competition math. Ran 35 hours autonomously across 1,158 tool calls and supports external harnesses like Claude Code.
94
Quality
$5.00
Blended/1M
90
tok/s
Grok 4.3
xAI's May 6 2026 flagship. 1M context, native video input, and a ~40% input price cut over Grok 4.20. Artificial Analysis Intelligence Index 53; outperforms Opus 4.7 ~1.26x on agentic Vending-Bench.
93
Quality
$1.88
Blended/1M
83
tok/s
Claude Sonnet 5
Anthropic's June 25 2026 mid-tier model and the workhorse of the Claude 5 family. It carries most of the Opus-line coding and agent quality at roughly a third of the price: SWE-bench Verified 82.4%, Terminal-Bench 2.1 68.1%, Artificial Analysis Intelligence Index 57.2. At $3/$15 (with a 90% caching discount) it is the default pick for high-volume agent work where Opus 4.8 or Fable 5 would be overkill on cost. 1M-token context and the same tool-use API as the flagship line.
93
Quality
$9.00
Blended/1M
98
tok/s
Gemini 2.5 Pro
Google's thinking model with native tool use, 1M context window, and strong multimodal capabilities.
92
Quality
$5.63
Blended/1M
87
tok/s
Kimi K2.6
Moonshot AI's April 20 2026 frontier model. 256K context with text, image, and video input. Artificial Analysis Intelligence Index 54; SWE-bench Verified 80.2%, up sharply from K2.5.
92
Quality
$2.11
Blended/1M
48
tok/s
DeepSeek V4.5
DeepSeek's July 5 2026 open-weight release under an MIT license — the strongest open reasoning model on price-per-quality. SWE-bench Pro 62.1%, GPQA Diamond 91.8%, and competition-math scores that sit with the proprietary frontier. Hosted API pricing is $0.50/$1.10 per 1M tokens; weights are downloadable for on-prem or private-cloud inference, so the marginal token cost self-hosted is infrastructure only. Benchmarks were self-reported at launch; the technical report and full eval harness followed within a week.
92
Quality
$0.80
Blended/1M
62
tok/s
Claude Opus 4
Anthropic's most capable model. Excels at complex analysis, nuanced writing, and extended agentic tasks.
91
Quality
$45.00
Blended/1M
52
tok/s
DeepSeek R1
DeepSeek's reasoning model. Competitive with o3 on math and coding at a fraction of the cost.
91
Quality
$1.37
Blended/1M
35
tok/s
Llama 5
Meta's June 30 2026 return to the frontier and the first Llama to compete with the proprietary top tier on general reasoning. Open weights under the Llama 5 Community License: MMLU-Pro 88.5, SWE-bench Verified 79.6%, strong multilingual coverage across 40+ languages, and a 1M-token context. Hosted at roughly $0.80/$2.40 on major providers, or self-hosted for infrastructure cost only. The natural open-weight default when you want frontier-adjacent quality with full deployment control and a permissive-enough license for most enterprise use.
91
Quality
$1.60
Blended/1M
76
tok/s
DeepSeek V4 Pro
1.6T MoE / 49B active. Apache 2.0, 1M context. Best price-per-quality of any frontier-tier model by a wide margin — regular rate $1.74/$3.48, recently offered at a discounted launch-promo rate near $0.44/$0.87.
90
Quality
$2.61
Blended/1M
33
tok/s
Claude Sonnet 4.6
Anthropic's Feb 17 2026 balanced model. Near-Opus performance at Sonnet pricing — SWE-bench Verified 79.6%, OSWorld 72.5%, 1M context. Best price-to-performance in the Claude lineup.
90
Quality
$9.00
Blended/1M
73
tok/s
GPT-4.1
OpenAI's latest flagship with 1M token context, improved instruction following and coding.
89
Quality
$5.00
Blended/1M
120
tok/s
MiniMax M3
MiniMax's June 1 2026 open-weight flagship — the first open model to combine frontier agentic coding, native multimodality, and a 1M-token context. Built on MiniMax Sparse Attention (MSA): ~1/20th the per-token compute at 1M context, >9x faster prefill, >15x faster decode. SWE-bench Pro 59.0% surpasses GPT-5.5 and Gemini 3.1 Pro and approaches Opus 4.7; Terminal-Bench 2.1 66.0%, OSWorld-Verified 70.06%. Benchmarks were self-reported and unverified at launch (weights and technical report due within ~10 days). Launch-promo pricing was $0.30/$1.20 per 1M tokens.
89
Quality
$1.50
Blended/1M
80
tok/s
o3 Mini
OpenAI's compact reasoning model with extended thinking capabilities for complex problem solving.
88
Quality
$2.75
Blended/1M
155
tok/s
Claude Sonnet 4
Anthropic's balanced model with excellent coding and reasoning. Best price-to-performance ratio.
88
Quality
$9.00
Blended/1M
95
tok/s
GLM-5.1
Z.ai's (Zhipu) 2026 flagship. 200K context, strong agentic and tool-use scores — 98% on τ²-Bench Telecom, on par with Grok 4.3. Artificial Analysis Intelligence Index 51.
88
Quality
$2.03
Blended/1M
48
tok/s
Grok 3
xAI's flagship model with strong reasoning and real-time information access. Trained on the Colossus cluster.
87
Quality
$9.00
Blended/1M
82
tok/s
DeepSeek V3
671B MoE model with 37B active parameters. Outstanding price-performance ratio and coding ability.
86
Quality
$0.69
Blended/1M
62
tok/s
Qwen 3.6 Plus
Alibaba's April 2026 flagship. Long-context multilingual specialist with strong CN/EN/AR coverage and competitive coding scores.
86
Quality
$3.50
Blended/1M
124
tok/s
GPT-4o
OpenAI's flagship multimodal model with vision, code generation, and function calling. Excellent all-round performance.
85
Quality
$6.25
Blended/1M
109
tok/s
Llama 4 Maverick
Meta's mixture-of-experts model with 17B active parameters and 128 experts. Strong multimodal and multilingual performance.
80
Quality
$0.40
Blended/1M
135
tok/s
Qwen 2.5 72B
Alibaba's flagship open-source model. Competitive with GPT-4o class models on benchmarks at a fraction of the cost.
80
Quality
$0.60
Blended/1M
85
tok/s
DeepSeek V4 Flash
284B MoE / 13B active. Apache 2.0, 1M context. Among the cheapest frontier-adjacent models — $0.10 input, $0.20 output per 1M tokens.
80
Quality
$0.15
Blended/1M
105
tok/s
Mistral Large 2
Mistral's flagship 123B model with strong multilingual and coding performance. Supports 128K context.
79
Quality
$4.00
Blended/1M
78
tok/s
Grok 3 Mini
xAI's efficient reasoning model with thinking capabilities at a lower cost point.
78
Quality
$0.40
Blended/1M
165
tok/s
Sonar Pro
Perplexity's search-augmented model. Combines LLM reasoning with real-time web search and citations.
78
Quality
$9.00
Blended/1M
65
tok/s
Codestral
Mistral's dedicated code model. Optimized for code generation, completion, and review across 80+ languages.
76
Quality
$0.60
Blended/1M
195
tok/s
Nemotron 3 Nano Omni
NVIDIA's 30B open multimodal model running vision + audio + text in a single stack. Tops 6 specialty leaderboards. April 2026.
76
Quality
$0.00
Blended/1M
158
tok/s
Claude 3.5 Haiku
Anthropic's fastest model. Ultra-low latency for real-time applications and high-volume tasks.
75
Quality
$2.40
Blended/1M
172
tok/s
Gemma 4 27B
Google's open-weight flagship under Apache 2.0 — April 2026. Designed for self-host with strong instruction-following and tool calling.
75
Quality
$0.00
Blended/1M
142
tok/s
Gemini 2.0 Flash
Google's fastest model. Optimized for speed and efficiency with strong coding and reasoning.
74
Quality
$0.25
Blended/1M
244
tok/s
Qwen 2.5 Coder 32B
Specialized coding model from Alibaba. Top open-source code model on HumanEval and SWE-Bench.
74
Quality
$0.30
Blended/1M
125
tok/s
GPT-4o Mini
Fast and affordable small model for lightweight tasks and high-throughput use cases.
72
Quality
$0.38
Blended/1M
183
tok/s
Llama 4 Scout
Meta's efficient MoE model with 16 experts. 10M token context window and strong multilingual support.
71
Quality
$0.28
Blended/1M
198
tok/s
Amazon Nova Pro
Amazon's capable multimodal model. Strong balance of accuracy, speed, and cost for diverse tasks.
70
Quality
$2.00
Blended/1M
110
tok/s
Command R+
Cohere's flagship for enterprise RAG. Optimized for retrieval-augmented generation and tool use.
68
Quality
$6.25
Blended/1M
72
tok/s