AI Model Directory

Detailed specs, pricing, and benchmarks for every major AI model.

Anthropic

Claude Fable 5

Anthropic's June 9 2026 flagship and the first generally available model in the Mythos class — the new #1 overall. State-of-the-art on nearly every tested benchmark: SWE-bench Pro 80.3% (vs Opus 4.8 69.2%, GPT-5.5 58.6%), FrontierCode Diamond 29.3%, OSWorld-Verified computer use 85.0%, and a leading GDPval-AA knowledge-work Elo of 1932. Routes high-risk cyber/bio prompts to Opus 4.8 as a safety fallback. $10/$50 pricing (2x Opus 4.8) with a 90% prompt-caching discount; exact context window and max output were not officially disclosed at launch.

100

Quality

$30.00

Blended/1M

58

tok/s

Best for: Frontier agentic coding & knowledge work
Anthropic

Claude Opus 4.8

Anthropic's May 28 2026 flagship and the new #1 on the Artificial Analysis Intelligence Index at 61.4 (+4.1 over Opus 4.7, +1.2 ahead of GPT-5.5). SWE-bench Verified 88.6%, SWE-bench Pro 69.2%, Terminal-Bench 2.1 74.6%, and a leading 49.8% on Humanity's Last Exam. Strongest computer-use/browser-agent model tested (Online-Mind2Web 84%). Same $5/$25 pricing as Opus 4.7; fast mode is ~2.5x faster and ~3x cheaper.

99

Quality

$15.00

Blended/1M

72

tok/s

Best for: Coding, agents & computer use
OpenAI

GPT-5.5 Pro

High-compute variant of GPT-5.5 with extended thinking. 6x the price for mid-teens AAII uplift on hard reasoning workloads.

98

Quality

$105.00

Blended/1M

68

tok/s

Best for: Reasoning at any cost
OpenAI
New

GPT-5.6

OpenAI's July 8 2026 flagship — a point release over GPT-5.5 that closes most of the gap to Claude Opus 4.8. Artificial Analysis Intelligence Index 61.0 (up from GPT-5.5's 60.2), SWE-bench Verified 88.1%, SWE-bench Pro 66.4%, Terminal-Bench 2.1 72.9%. The headline change is agentic reliability: OpenAI reports far fewer dropped tool calls on long multi-step runs. Pricing is unchanged from GPT-5.5 at $5/$30 with a 90% cached-input discount, so the quality gain is effectively free for existing users.

98

Quality

$17.50

Blended/1M

96

tok/s

Best for: General frontier reasoning & tools
OpenAI

GPT-5.5

OpenAI's "Spud" — the first fully retrained base model since GPT-4.5. 1M context window; tops the Artificial Analysis Intelligence Index at 59-60 and the LMArena text leaderboard at ~1506 Elo. Shipped April 23 2026, on AWS Bedrock April 28.

97

Quality

$17.50

Blended/1M

70

tok/s

Best for: Frontier general purpose
Google
New

Gemini 3.2 Pro

Google's July 2 2026 flagship and a clear step over Gemini 3.1 Pro. Artificial Analysis Intelligence Index 59.6, SWE-bench Verified 86.4%, and a class-leading 2M-token context that holds accuracy deeper into the window than any rival on Google's own long-context evals. Still the value pick among frontier proprietary models at $2/$12 (unchanged from 3.1 Pro), with native audio and video understanding. Best all-round choice when the workload is long documents, multimodal input, or high request volume where price matters.

97

Quality

$7.00

Blended/1M

128

tok/s

Best for: Long-context & multimodal
Anthropic

Claude Opus 4.7

Anthropic's April 16 2026 flagship. SWE-bench Verified 87.6%, SWE-bench Pro 64.3%, GPQA Diamond 94.2%. Sits in the LMArena top tier at ~1503 (1505 thinking). Same $5/$25 pricing as Opus 4.6.

96

Quality

$15.00

Blended/1M

68

tok/s

Best for: Coding & agentic workflows
Google

Gemini 3.1 Pro

Tied at the top of the LMArena text leaderboard May 2026 @ ~1505 Elo. GPQA Diamond 94.3%, MMLU-Pro 91.0% — current scientific reasoning leader. 1M token context; fastest frontier model at ~131 tok/s.

96

Quality

$7.00

Blended/1M

131

tok/s

Best for: Science & long-context
OpenAI

o3

OpenAI's most powerful reasoning model. State-of-the-art on MATH, coding, and science benchmarks.

94

Quality

$25.00

Blended/1M

68

tok/s

Best for: Hard reasoning
Alibaba Cloud

Qwen 3.7 Max

Alibaba's May 20 2026 proprietary flagship, unveiled at the Alibaba Cloud Summit in Hangzhou. Highest-ranked Chinese model on the Artificial Analysis Intelligence Index at 56.6 (#5 overall, +4.8 over Qwen 3.6 Max Preview). 1M context; SWE-bench Pro 60.6, Terminal-Bench 2.0 69.7, GPQA Diamond 92.4, and a table-leading 97.1 on HMMT Feb 2026 competition math. Ran 35 hours autonomously across 1,158 tool calls and supports external harnesses like Claude Code.

94

Quality

$5.00

Blended/1M

90

tok/s

Best for: Long autonomous agentic runs
xAI

Grok 4.3

xAI's May 6 2026 flagship. 1M context, native video input, and a ~40% input price cut over Grok 4.20. Artificial Analysis Intelligence Index 53; outperforms Opus 4.7 ~1.26x on agentic Vending-Bench.

93

Quality

$1.88

Blended/1M

83

tok/s

Best for: Agentic tasks & real-time info
Anthropic
New

Claude Sonnet 5

Anthropic's June 25 2026 mid-tier model and the workhorse of the Claude 5 family. It carries most of the Opus-line coding and agent quality at roughly a third of the price: SWE-bench Verified 82.4%, Terminal-Bench 2.1 68.1%, Artificial Analysis Intelligence Index 57.2. At $3/$15 (with a 90% caching discount) it is the default pick for high-volume agent work where Opus 4.8 or Fable 5 would be overkill on cost. 1M-token context and the same tool-use API as the flagship line.

93

Quality

$9.00

Blended/1M

98

tok/s

Best for: Balanced agents & coding value
Google

Gemini 2.5 Pro

Google's thinking model with native tool use, 1M context window, and strong multimodal capabilities.

92

Quality

$5.63

Blended/1M

87

tok/s

Best for: Multimodal + value
Moonshot AI

Kimi K2.6

Moonshot AI's April 20 2026 frontier model. 256K context with text, image, and video input. Artificial Analysis Intelligence Index 54; SWE-bench Verified 80.2%, up sharply from K2.5.

92

Quality

$2.11

Blended/1M

48

tok/s

Best for: Frontier quality at low cost
DeepSeek
NewOSS

DeepSeek V4.5

DeepSeek's July 5 2026 open-weight release under an MIT license — the strongest open reasoning model on price-per-quality. SWE-bench Pro 62.1%, GPQA Diamond 91.8%, and competition-math scores that sit with the proprietary frontier. Hosted API pricing is $0.50/$1.10 per 1M tokens; weights are downloadable for on-prem or private-cloud inference, so the marginal token cost self-hosted is infrastructure only. Benchmarks were self-reported at launch; the technical report and full eval harness followed within a week.

92

Quality

$0.80

Blended/1M

62

tok/s

Best for: Open-weight reasoning value
Anthropic

Claude Opus 4

Anthropic's most capable model. Excels at complex analysis, nuanced writing, and extended agentic tasks.

91

Quality

$45.00

Blended/1M

52

tok/s

Best for: Complex analysis
DeepSeek
OSS

DeepSeek R1

DeepSeek's reasoning model. Competitive with o3 on math and coding at a fraction of the cost.

91

Quality

$1.37

Blended/1M

35

tok/s

Best for: Cheap reasoning
Meta
NewOSS

Llama 5

Meta's June 30 2026 return to the frontier and the first Llama to compete with the proprietary top tier on general reasoning. Open weights under the Llama 5 Community License: MMLU-Pro 88.5, SWE-bench Verified 79.6%, strong multilingual coverage across 40+ languages, and a 1M-token context. Hosted at roughly $0.80/$2.40 on major providers, or self-hosted for infrastructure cost only. The natural open-weight default when you want frontier-adjacent quality with full deployment control and a permissive-enough license for most enterprise use.

91

Quality

$1.60

Blended/1M

76

tok/s

Best for: Open-weight general & multilingual
DeepSeek
OSS

DeepSeek V4 Pro

1.6T MoE / 49B active. Apache 2.0, 1M context. Best price-per-quality of any frontier-tier model by a wide margin — regular rate $1.74/$3.48, recently offered at a discounted launch-promo rate near $0.44/$0.87.

90

Quality

$2.61

Blended/1M

33

tok/s

Best for: Open-source value leader
Anthropic

Claude Sonnet 4.6

Anthropic's Feb 17 2026 balanced model. Near-Opus performance at Sonnet pricing — SWE-bench Verified 79.6%, OSWorld 72.5%, 1M context. Best price-to-performance in the Claude lineup.

90

Quality

$9.00

Blended/1M

73

tok/s

Best for: Coding & balance
OpenAI

GPT-4.1

OpenAI's latest flagship with 1M token context, improved instruction following and coding.

89

Quality

$5.00

Blended/1M

120

tok/s

Best for: Long context
MiniMax
OSS

MiniMax M3

MiniMax's June 1 2026 open-weight flagship — the first open model to combine frontier agentic coding, native multimodality, and a 1M-token context. Built on MiniMax Sparse Attention (MSA): ~1/20th the per-token compute at 1M context, >9x faster prefill, >15x faster decode. SWE-bench Pro 59.0% surpasses GPT-5.5 and Gemini 3.1 Pro and approaches Opus 4.7; Terminal-Bench 2.1 66.0%, OSWorld-Verified 70.06%. Benchmarks were self-reported and unverified at launch (weights and technical report due within ~10 days). Launch-promo pricing was $0.30/$1.20 per 1M tokens.

89

Quality

$1.50

Blended/1M

80

tok/s

Best for: Open-weight agentic coding
OpenAI

o3 Mini

OpenAI's compact reasoning model with extended thinking capabilities for complex problem solving.

88

Quality

$2.75

Blended/1M

155

tok/s

Best for: Reasoning & math
Anthropic

Claude Sonnet 4

Anthropic's balanced model with excellent coding and reasoning. Best price-to-performance ratio.

88

Quality

$9.00

Blended/1M

95

tok/s

Best for: Coding & balance
Z.ai (Zhipu AI)
OSS

GLM-5.1

Z.ai's (Zhipu) 2026 flagship. 200K context, strong agentic and tool-use scores — 98% on τ²-Bench Telecom, on par with Grok 4.3. Artificial Analysis Intelligence Index 51.

88

Quality

$2.03

Blended/1M

48

tok/s

Best for: Open-weight agentic & tool use
xAI

Grok 3

xAI's flagship model with strong reasoning and real-time information access. Trained on the Colossus cluster.

87

Quality

$9.00

Blended/1M

82

tok/s

Best for: Real-time info
DeepSeek
OSS

DeepSeek V3

671B MoE model with 37B active parameters. Outstanding price-performance ratio and coding ability.

86

Quality

$0.69

Blended/1M

62

tok/s

Best for: Best open-source value
Alibaba Cloud

Qwen 3.6 Plus

Alibaba's April 2026 flagship. Long-context multilingual specialist with strong CN/EN/AR coverage and competitive coding scores.

86

Quality

$3.50

Blended/1M

124

tok/s

Best for: Multilingual & APAC
OpenAI

GPT-4o

OpenAI's flagship multimodal model with vision, code generation, and function calling. Excellent all-round performance.

85

Quality

$6.25

Blended/1M

109

tok/s

Best for: General purpose
Meta
OSS

Llama 4 Maverick

Meta's mixture-of-experts model with 17B active parameters and 128 experts. Strong multimodal and multilingual performance.

80

Quality

$0.40

Blended/1M

135

tok/s

Best for: Open-source value
Alibaba Cloud
OSS

Qwen 2.5 72B

Alibaba's flagship open-source model. Competitive with GPT-4o class models on benchmarks at a fraction of the cost.

80

Quality

$0.60

Blended/1M

85

tok/s

Best for: Open-source flagship
DeepSeek
OSS

DeepSeek V4 Flash

284B MoE / 13B active. Apache 2.0, 1M context. Among the cheapest frontier-adjacent models — $0.10 input, $0.20 output per 1M tokens.

80

Quality

$0.15

Blended/1M

105

tok/s

Best for: Cheap-and-fast cascade tier
Mistral AI

Mistral Large 2

Mistral's flagship 123B model with strong multilingual and coding performance. Supports 128K context.

79

Quality

$4.00

Blended/1M

78

tok/s

Best for: Multilingual
xAI

Grok 3 Mini

xAI's efficient reasoning model with thinking capabilities at a lower cost point.

78

Quality

$0.40

Blended/1M

165

tok/s

Best for: Budget reasoning
Perplexity

Sonar Pro

Perplexity's search-augmented model. Combines LLM reasoning with real-time web search and citations.

78

Quality

$9.00

Blended/1M

65

tok/s

Best for: Search + citations
Mistral AI

Codestral

Mistral's dedicated code model. Optimized for code generation, completion, and review across 80+ languages.

76

Quality

$0.60

Blended/1M

195

tok/s

Best for: Code generation
Mistral AI
OSS

Nemotron 3 Nano Omni

NVIDIA's 30B open multimodal model running vision + audio + text in a single stack. Tops 6 specialty leaderboards. April 2026.

76

Quality

$0.00

Blended/1M

158

tok/s

Best for: Open multimodal
Anthropic

Claude 3.5 Haiku

Anthropic's fastest model. Ultra-low latency for real-time applications and high-volume tasks.

75

Quality

$2.40

Blended/1M

172

tok/s

Best for: Speed & cost
Google
OSS

Gemma 4 27B

Google's open-weight flagship under Apache 2.0 — April 2026. Designed for self-host with strong instruction-following and tool calling.

75

Quality

$0.00

Blended/1M

142

tok/s

Best for: Self-hosted general purpose
Google

Gemini 2.0 Flash

Google's fastest model. Optimized for speed and efficiency with strong coding and reasoning.

74

Quality

$0.25

Blended/1M

244

tok/s

Best for: Fastest + cheapest
Alibaba Cloud
OSS

Qwen 2.5 Coder 32B

Specialized coding model from Alibaba. Top open-source code model on HumanEval and SWE-Bench.

74

Quality

$0.30

Blended/1M

125

tok/s

Best for: Open-source coding
OpenAI

GPT-4o Mini

Fast and affordable small model for lightweight tasks and high-throughput use cases.

72

Quality

$0.38

Blended/1M

183

tok/s

Best for: High throughput
Meta
OSS

Llama 4 Scout

Meta's efficient MoE model with 16 experts. 10M token context window and strong multilingual support.

71

Quality

$0.28

Blended/1M

198

tok/s

Best for: Longest context
Amazon

Amazon Nova Pro

Amazon's capable multimodal model. Strong balance of accuracy, speed, and cost for diverse tasks.

70

Quality

$2.00

Blended/1M

110

tok/s

Best for: AWS ecosystem
Cohere

Command R+

Cohere's flagship for enterprise RAG. Optimized for retrieval-augmented generation and tool use.

68

Quality

$6.25

Blended/1M

72

tok/s

Best for: Enterprise RAG