AI Model Directory
Detailed specs, pricing, and benchmarks for every major AI model.
Claude Fable 5
Anthropic's June 9 2026 flagship and the first generally available model in the Mythos class: the new #1 overall. State-of-the-art on nearly every tested benchmark: SWE-bench Pro 80.3% (vs Opus 4.8 69.2%, GPT-5.5 58.6%), FrontierCode Diamond 29.3%, OSWorld-Verified computer use 85.0%, and a leading GDPval-AA knowledge-work Elo of 1932. Routes high-risk cyber/bio prompts to Opus 4.8 as a safety fallback. $10/$50 pricing (2x Opus 4.8) with a 90% prompt-caching discount; exact context window and max output were not officially disclosed at launch.
100
Quality
$30.00
Blended/1M
58
tok/s
Claude Mythos 5
The top-scoring model on the board, and the one almost nobody can call: Mythos 5 is available only through Anthropic's Project Glasswing, succeeding the invitation-only Mythos Preview. Capabilities, pricing, limits, and API behaviour are identical to Claude Fable 5 — $10/$50 per 1M tokens, 1M context, 128K output, always-on thinking with the raw chain of thought never returned, and a 30-day data-retention requirement that rules out zero-retention orgs. Listed here for completeness of the frontier; if you are not in Glasswing, Claude Opus 5 or Fable 5 is the model you can actually buy.
100
Quality
$30.00
Blended/1M
56
tok/s
Claude Opus 5
Anthropic's 24 Jul 2026 flagship and the highest-scoring verified model on the current board. A step change over Opus 4.8 on deep reasoning, long-horizon agentic work, and test-time compute scaling — at unchanged Opus pricing of $5/$25 per 1M tokens, roughly half the cost of Claude Fable 5. 1M context (default and maximum), 128K max output. Thinking is on by default rather than opt-in, the prompt-cache minimum drops to 512 tokens from 1024, and the full low-through-max effort ladder is available; low and medium effort are unusually strong here, which makes it cheaper to run than the sticker price suggests. Fast mode is Claude-API-only at $10/$50. Draws on a separate rate-limit pool from the Opus 4.x models.
99
Quality
$15.00
Blended/1M
74
tok/s
Claude Opus 4.8
Anthropic's May 28 2026 flagship and the new #1 on the Artificial Analysis Intelligence Index at 61.4 (+4.1 over Opus 4.7, +1.2 ahead of GPT-5.5). SWE-bench Verified 88.6%, SWE-bench Pro 69.2%, Terminal-Bench 2.1 74.6%, and a leading 49.8% on Humanity's Last Exam. Strongest computer-use/browser-agent model tested (Online-Mind2Web 84%). Same $5/$25 pricing as Opus 4.7; fast mode is ~2.5x faster and ~3x cheaper.
98
Quality
$15.00
Blended/1M
72
tok/s
GPT-5.6 Sol
OpenAI's flagship tier, GA 9 Jul 2026. The 5.6 generation replaced the old one-model-with-a-dial approach with three durable tiers — Sol (flagship), Terra (balanced), Luna (cheap) — and `gpt-5.6` is now an alias for `gpt-5.6-sol`. There is no separate Pro model ID; Pro is a reasoning mode on Sol. Artificial Analysis Intelligence Index 61.0, SWE-bench Verified 88.1%, SWE-bench Pro 66.4%, Terminal-Bench 2.1 72.9%. 1.05M context (922K input / 128K output). Pricing holds at $5/$30 with cache reads at 10% of input; requests over 272K tokens move to a long-context meter at $10/$45.
98
Quality
$17.50
Blended/1M
96
tok/s
GPT-5.5
OpenAI's "Spud": the first fully retrained base model since GPT-4.5. 1M context window; tops the Artificial Analysis Intelligence Index at 59-60 and the LMArena text leaderboard at ~1506 Elo. Shipped April 23 2026, on AWS Bedrock April 28.
97
Quality
$17.50
Blended/1M
70
tok/s
Kimi K3
Moonshot AI's 16 Jul 2026 flagship and the clearest evidence that the open-weight tier has caught the frontier: a 2.8-trillion-parameter MoE (16 of 896 experts routed per token, Stable LatentMoE, MXFP4 weights) with a 1,048,576-token context and native text, vision, and video input. Artificial Analysis Intelligence Index 57, ranking #3 overall behind only Claude Fable 5 and GPT-5.6 Sol, and ahead of every other proprietary model on the board. It beats GLM-5.2 across Moonshot's harness — DeepSWE 67.5 vs 46.2, FrontierSWE 81.2 vs 67.3, SWE Marathon 42.0 vs 13.0, Terminal-Bench 2.1 88.3 vs 82.7, GPQA-Diamond 93.5 vs 91.2 — and tops the open board on Humanity's Last Exam (56%) and BrowseComp (91.2). At $3/$15 it undercuts both Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30) on input and output while scoring higher on independent benchmarks. Two caveats: K3 is a heavy token consumer, and it is a genuine cost outlier among open models at ~$0.94 per task versus DeepSeek V4 Pro's $0.04. Weights shipped on Hugging Face at moonshotai/Kimi-K3 (2.8T total, 104B active) under Moonshot's own Kimi K3 License rather than a standard open licence, so read the terms before building on it. Hallucination rate rose to 51% from K2.6's 39% — cross-check factual output.
97
Quality
$9.00
Blended/1M
55
tok/s
GPT-5.5 Pro
High-compute variant of GPT-5.5 with extended thinking. 6x the price for mid-teens AAII uplift on hard reasoning workloads.
96
Quality
$105.00
Blended/1M
68
tok/s
Claude Opus 4.7
Anthropic's April 16 2026 flagship. SWE-bench Verified 87.6%, SWE-bench Pro 64.3%, GPQA Diamond 94.2%. Sits in the LMArena top tier at ~1503 (1505 thinking). Same $5/$25 pricing as Opus 4.6.
96
Quality
$15.00
Blended/1M
68
tok/s
Gemini 3.1 Pro
Tied at the top of the LMArena text leaderboard May 2026 @ ~1505 Elo. GPQA Diamond 94.3%, MMLU-Pro 91.0%: current scientific reasoning leader. 1M token context; fastest frontier model at ~131 tok/s.
96
Quality
$7.00
Blended/1M
131
tok/s
Qwen3.8 Max
Reached GA on 3 Aug 2026 and, around 12–14 Aug, became the first Qwen-Max-class model Alibaba has ever open-weighted — following its 19 Jul preview at WAIC Shanghai. Still a 2.4T-parameter sparse MoE (active-parameter count undisclosed), but treat the hosted API and the open weights as two different products: the published checkpoint is reportedly text-only, dropping the vision and 1M-context features the hosted API keeps. Weights ship under a new custom revenue-share license, not Apache 2.0 — large commercial 'model-as-a-service' deployments above an unspecified threshold need a separate agreement with Alibaba, so check the license before redistributing. Vendor-reported scores at GA: Terminal-Bench 2.1 86.6 (vs. Claude Opus 4.8/Fable 5 at 84.6, GPT-5.6 Sol at 88.8), SWE-bench Pro 67.7, GPQA Diamond 92.6, DeepSWE 1.1 56.6 (up from 21.6 on Qwen3.7 Max) — independent Arena/AA verification is still catching up as of 17 Aug 2026, though early Arena placement (~1491) already makes it the highest-ranked open-weight model on the board. Hosted API pricing dropped from the preview placeholder to $2/$6 per 1M tokens ($0.25 cached).
96
Quality
$4.00
Blended/1M
86
tok/s
o3
OpenAI's most powerful reasoning model. State-of-the-art on MATH, coding, and science benchmarks.
94
Quality
$5.00
Blended/1M
68
tok/s
Qwen 3.7 Max
Alibaba's May 20 2026 proprietary flagship, unveiled at the Alibaba Cloud Summit in Hangzhou. Highest-ranked Chinese model on the Artificial Analysis Intelligence Index at 56.6 (#5 overall, +4.8 over Qwen 3.6 Max Preview). 1M context; SWE-bench Pro 60.6, Terminal-Bench 2.0 69.7, GPQA Diamond 92.4, and a table-leading 97.1 on HMMT Feb 2026 competition math. Ran 35 hours autonomously across 1,158 tool calls and supports external harnesses like Claude Code.
94
Quality
$5.00
Blended/1M
90
tok/s
Grok 4.5
xAI's July 2026 flagship at $2 input / $6 output per 1M tokens — notable for keeping a 3:1 output-to-input ratio where most rivals charge 5-6x, which makes it materially cheaper than headline comparisons suggest on generation-heavy work. 500K context, Artificial Analysis Intelligence Index 54. Watch three things: prompts over 200K tokens double in price, Grok 4.5 is excluded from the batch discount that applies to Grok 4.3, and tools bill separately (Web/X Search and Code Execution at $5 per 1,000 calls, File Attachments at $10 per 1,000). Cached input is quoted between $0.30 and $0.50 depending on source — verify in the console before modelling costs.
94
Quality
$4.00
Blended/1M
88
tok/s
Grok 4.3
xAI's May 6 2026 flagship. 1M context, native video input, and a ~40% input price cut over Grok 4.20. Artificial Analysis Intelligence Index 53; outperforms Opus 4.7 ~1.26x on agentic Vending-Bench.
93
Quality
$1.88
Blended/1M
83
tok/s
GPT-5.6 Terra
The balanced everyday tier of the GPT-5.6 family, GA 9 Jul 2026. On 30 Jul 2026 OpenAI cut Terra 20% to $2 input / $12 output per 1M tokens, crediting inference work that reduced end-to-end serving cost by 20% and improved token-generation efficiency by more than 15%. Same 1.05M context as Sol, cache reads at $0.20, Batch API at half rate. This is the tier most production workloads should default to — it lands within a few points of Sol on most evals at 40% of the price.
93
Quality
$7.00
Blended/1M
118
tok/s
Claude Sonnet 5
Anthropic's 25 Jun 2026 mid-tier model and the workhorse of the Claude 5 family. It carries most of the Opus-line coding and agent quality at roughly a third of the price: SWE-bench Verified 82.4%, Terminal-Bench 2.1 68.1%, Artificial Analysis Intelligence Index 57.2. List price is $3/$15 with a 90% caching discount, but an introductory rate of $2/$10 runs through 31 Aug 2026 — worth locking in evaluation work before it lapses. 1M-token context and the same tool-use API as the flagship line. The default pick for high-volume agent work where Opus 5 or Fable 5 would be overkill on cost.
93
Quality
$9.00
Blended/1M
98
tok/s
Gemini 2.5 Pro
Google's thinking model with native tool use, 1M context window, and strong multimodal capabilities.
92
Quality
$5.63
Blended/1M
87
tok/s
Kimi K2.6
Moonshot AI's April 20 2026 frontier model. 256K context with text, image, and video input. Artificial Analysis Intelligence Index 54; SWE-bench Verified 80.2%, up sharply from K2.5.
92
Quality
$2.11
Blended/1M
48
tok/s
Claude Opus 4
Anthropic's most capable model. Excels at complex analysis, nuanced writing, and extended agentic tasks.
91
Quality
$45.00
Blended/1M
52
tok/s
DeepSeek R1
DeepSeek's reasoning model. Competitive with o3 on math and coding at a fraction of the cost.
91
Quality
$1.37
Blended/1M
35
tok/s
GLM-5.2
Z.ai's 13 Jun 2026 open-weight flagship, and the model that held the top open-weight slot until Kimi K3 landed a month later. 744B MoE / ~40B active, MIT-licensed weights on Hugging Face, 1M context, 131K max output. $1.40 input / $4.40 output per 1M tokens ($0.26 cached) — a blended ~$0.90/1M. The headline result is SWE-bench Pro 62.1%, which beats GPT-5.5's 58.6%: an MIT-licensed model you can self-host outscoring a $5/$30 US flagship on agentic coding. Also the fastest of the big open MoEs at ~168 tokens/sec, roughly 2.7x Kimi K3 and DeepSeek V4 Pro. Self-hosting needs about 1TB VRAM in BF16, or ~8x H200 at FP8.
91
Quality
$2.90
Blended/1M
168
tok/s
Claude Sonnet 4.6
Anthropic's Feb 17 2026 balanced model. Near-Opus performance at Sonnet pricing: SWE-bench Verified 79.6%, OSWorld 72.5%, 1M context. Best price-to-performance in the Claude lineup.
90
Quality
$9.00
Blended/1M
73
tok/s
Gemini 3.6 Flash
Google's 21 Jul 2026 mid-tier release, replacing Gemini 3.5 Flash. Introductory pricing of $0.75 input / $3.75 output per 1M tokens runs through 31 Dec 2026, after which the standard $1.50 / $7.50 rate applies; cached input is $0.15 and batch mode is half rate. 1M context, native audio and video understanding, and the fastest response times in the frontier-adjacent band. Note that Google shipped no Gemini 3.2–3.5 Pro: the Pro line still tops out at Gemini 3.1 Pro, and the 3.5/3.6 releases are all Flash-tier.
90
Quality
$2.25
Blended/1M
148
tok/s
GPT-4.1
OpenAI's latest flagship with 1M token context, improved instruction following and coding.
89
Quality
$5.00
Blended/1M
120
tok/s
DeepSeek V4 Pro
Left preview and reached GA on 12–13 Aug 2026 as the '0813' checkpoint — same 1.6T MoE / 49B-active architecture as the April launch, but re-post-trained and, notably, repriced upward. The promotional $0.435/$0.87 rate is gone: GA pricing is $0.66 input / $1.98 output per 1M tokens, a roughly 52% input and 128% output increase, which DeepSeek framed as stepping back from a pure race-to-zero strategy. Still MIT-licensed weights on Hugging Face, 1M context. Vendor-reported scores are strong — LiveCodeBench 93.5%, SWE-bench Verified 80.6% (matching Gemini 3.1 Pro), GPQA Diamond 90.1%, Codeforces 3206 — but treat them cautiously: contamination-resistant independent benchmarks (DeepSWE) and Code Arena both show a wider gap to closed frontier models than DeepSeek's own numbers suggest, and third-party Arena Elo for the 0813 build is still preliminary (~1458–1465) as of 17 Aug 2026.
89
Quality
$1.32
Blended/1M
62
tok/s
MiniMax M3
MiniMax's 1 Jun 2026 open-weight flagship: the first open model to combine frontier agentic coding, native multimodality, and a 1M-token context. Built on MiniMax Sparse Attention (MSA): ~1/20th the per-token compute at 1M context, >9x faster prefill, >15x faster decode. Now third-party verified and holding up well — SWE-bench 80.5% (second among open models, behind only DeepSeek V4 Pro's 80.6%), GPQA Diamond 93.0% (second behind Kimi K3), OSWorld 70.1%, BrowseComp 83.5%, Terminal-Bench 2.1 66.0%. At $0.60/$2.40 it delivers SWE-bench parity with models costing 10x more. Launch-promo pricing was $0.30/$1.20.
89
Quality
$1.50
Blended/1M
80
tok/s
o3 Mini
OpenAI's compact reasoning model with extended thinking capabilities for complex problem solving.
88
Quality
$2.75
Blended/1M
155
tok/s
Claude Sonnet 4
Anthropic's balanced model with excellent coding and reasoning. Best price-to-performance ratio.
88
Quality
$9.00
Blended/1M
95
tok/s
GLM-5.1
Z.ai's (Zhipu) 2026 flagship. 200K context, strong agentic and tool-use scores: 98% on τ²-Bench Telecom, on par with Grok 4.3. Artificial Analysis Intelligence Index 51.
88
Quality
$2.03
Blended/1M
48
tok/s
Grok 3
xAI's flagship model with strong reasoning and real-time information access. Trained on the Colossus cluster.
87
Quality
$9.00
Blended/1M
82
tok/s
DeepSeek V3
671B MoE model with 37B active parameters. Outstanding price-performance ratio and coding ability.
86
Quality
$0.69
Blended/1M
62
tok/s
Qwen 3.6 Plus
Alibaba's April 2026 flagship. Long-context multilingual specialist with strong CN/EN/AR coverage and competitive coding scores.
86
Quality
$3.50
Blended/1M
124
tok/s
GPT-4o
OpenAI's flagship multimodal model with vision, code generation, and function calling. Excellent all-round performance.
85
Quality
$6.25
Blended/1M
109
tok/s
DeepSeek V4 Flash
The cheapest usable model in the directory at $0.14 input / $0.28 output per 1M tokens, with cache hits at $0.0028 — DeepSeek cut the cache-hit rate to a tenth of its launch price on 26 Apr 2026. 284B MoE / 13B active, MIT weights, 1M context, and an OpenAI-compatible endpoint. On 31 Jul 2026 the `deepseek-v4-flash` API ID began serving the V4-Flash-0731 checkpoint, which posts 79% on SWE-bench Verified and 85.9 on BrowseComp — numbers that would have been frontier a year earlier, at roughly 1/100th of GPT-5.6 Sol's output price. The hosted service is in public beta. Rate limit is 2,500 concurrent requests.
84
Quality
$0.21
Blended/1M
105
tok/s
GPT-5.6 Luna
The fastest, cheapest member of the GPT-5.6 family and the site of the steepest price cut of the year from a US lab: on 30 Jul 2026 OpenAI dropped Luna 80%, from $1/$6 to $0.20 input / $1.20 output per 1M tokens. Cache reads run $0.02. That is an explicit answer to DeepSeek and the Chinese open-weight tier on price — though at $0.20/$1.20 Luna is still comfortably more expensive than DeepSeek V4 Flash at $0.14/$0.28. Full 1.05M context, unlike most rivals' cheap tiers.
83
Quality
$0.70
Blended/1M
186
tok/s
Llama 4 Maverick
Meta's mixture-of-experts model with 17B active parameters and 128 experts. Strong multimodal and multilingual performance.
80
Quality
$0.40
Blended/1M
135
tok/s
Qwen 2.5 72B
Alibaba's flagship open-source model. Competitive with GPT-4o class models on benchmarks at a fraction of the cost.
80
Quality
$0.60
Blended/1M
85
tok/s
Qwen3.8 27B
Alibaba's fully open, Apache 2.0-licensed sibling to Qwen3.8 Max, released 13–14 Aug 2026 — the locally-runnable pick where the Max tier's custom revenue-share license doesn't apply. 27.8B dense params (not MoE) with native vision-language input and Hybrid Gated DeltaNet attention (3-of-4 sublayers linear attention), giving a native 262K context extensible to 1M via YaRN RoPE scaling. Vendor-reported agentic-coding gains over its predecessor are large: SWE-bench Pro 61.7, DeepSWE 1.1 42.2 (vs. 13.3 on Qwen3.6 27B), Terminal-Bench 2.1 73.0, LiveCodeBench v6 90.3, OSWorld 84.3 (vs. 63.9) — Alibaba claims it beats Claude Opus 4.6 Max on several of these. Runs in roughly 17GB VRAM at 4-bit, the main reason to reach for it over Max for self-hosting; one independent review found it markedly slower and more token-hungry per task than its predecessor at inference, so budget for that trade-off. No independent MMLU/Arena placement yet as of 17 Aug 2026.
80
Quality
$0.00
Blended/1M
—
tok/s
Mistral Large 2
Mistral's flagship 123B model with strong multilingual and coding performance. Supports 128K context.
79
Quality
$4.00
Blended/1M
78
tok/s
Gemini 3.5 Flash-Lite
Google's cheapest current-generation model at $0.30 input / $2.50 output per 1M tokens ($0.03 cached), released 21 Jul 2026 alongside 3.6 Flash. 1M context. Competitive on high-volume classification and extraction work, though DeepSeek V4 Flash still undercuts it roughly 2x on input and 9x on output.
79
Quality
$1.40
Blended/1M
192
tok/s
Nemotron 3.5 Lightning
NVIDIA's 11 Aug 2026 release, the smallest model in the Nemotron 3 family (sibling to Nemotron 3 Ultra, distinct from the catalog's Nemotron 3 Nano Omni). Hybrid Mamba-2 + MoE + attention architecture, 30B total / 3B active params, pretrained on 20T+ tokens with an NVFP4 recipe, under NVIDIA's permissive OpenMDW-1.1 license (open weights, training data, and recipes; commercial use allowed). 1M-token context. Ships with multi-token-prediction draft models (DSpark, DFlash) for speculative decoding and is single-GPU deployable (1x DGX Spark or 1x H100). Alongside it NVIDIA released NeMo Switchyard, an open-source agent-routing library. Vendor-reported: MMLU Pro 81.9, GPQA Diamond 75.4, SWE-bench Verified 51.6-52.8, PinchBench 86% accuracy (30% faster than Qwen3.6 35B). No independent Arena placement yet as of 17 Aug 2026; hosted via build.nvidia.com and OpenRouter, but paid-tier pricing wasn't confirmed at the time of writing — verify before quoting a rate.
79
Quality
$0.00
Blended/1M
—
tok/s
Grok 3 Mini
xAI's efficient reasoning model with thinking capabilities at a lower cost point.
78
Quality
$0.40
Blended/1M
165
tok/s
Sonar Pro
Perplexity's search-augmented model. Combines LLM reasoning with real-time web search and citations.
78
Quality
$9.00
Blended/1M
65
tok/s
Codestral
Mistral's dedicated code model. Optimized for code generation, completion, and review across 80+ languages.
76
Quality
$0.60
Blended/1M
195
tok/s
Nemotron 3 Nano Omni
NVIDIA's 30B open multimodal model running vision + audio + text in a single stack. Tops 6 specialty leaderboards. April 2026.
76
Quality
$0.00
Blended/1M
158
tok/s
Hunyuan Hy3
Tencent's 6 Jul 2026 full release (previewed 23 Apr 2026), a 295B-total / 21B-active MoE plus a separate 3.8B multi-token-prediction layer, under Apache 2.0 with no field-of-use or geographic restrictions. 256K native context. Tencent concedes the agentic-coding crown to GLM-5.2 — SWE-bench Pro rose from 46.0 at preview to 57.9 at release, still behind GLM — but claims leads in agentic search, tool use, and long-context retrieval at under half GLM-5.2's memory footprint. First Tencent entry in this catalog; Hunyuan's flagship chat line (outside this Hy3 open release) remains closed-weight/API-only. No independent Arena placement yet as of 17 Aug 2026.
76
Quality
$0.00
Blended/1M
—
tok/s
Claude 3.5 Haiku
Anthropic's fastest model. Ultra-low latency for real-time applications and high-volume tasks.
75
Quality
$2.40
Blended/1M
172
tok/s
Gemma 4 27B
Google's open-weight flagship under Apache 2.0: April 2026. Designed for self-host with strong instruction-following and tool calling.
75
Quality
$0.00
Blended/1M
142
tok/s
Gemini 2.0 Flash
Google's fastest model. Optimized for speed and efficiency with strong coding and reasoning.
74
Quality
$0.25
Blended/1M
244
tok/s
Qwen 2.5 Coder 32B
Specialized coding model from Alibaba. Top open-source code model on HumanEval and SWE-Bench.
74
Quality
$0.30
Blended/1M
125
tok/s
Ling-3.0-Flash
First catalog entry for Ant Group's inclusionAI/Bailing team, an actively-shipping open-weight lab (Ling non-thinking, Ring thinking, Ming multimodal families) not previously tracked here despite multiple trillion-parameter MIT-licensed releases through 2026 (Ling-2.5-1T, Ring-2.5-1T, Ling-2.6-1T). Ling-3.0-Flash, announced 27 Jul 2026 with weights open-sourced 5 Aug 2026, is a 124B-total MoE with roughly 5-7B active params per token (sources give both figures; verify against the official Hugging Face card before treating either as final), 256K context, MIT license. Ant claims it matches or beats their own 1T-parameter Ling-2.6-1T flagship on most benchmarks at roughly 1/8th the total parameters — a notable efficiency claim, but independent verification and a standard benchmark table weren't available at the time of writing. Weights ship in BF16 (255GB) and FP8 (128GB) on Hugging Face and ModelScope.
74
Quality
$0.00
Blended/1M
—
tok/s
GPT-4o Mini
Fast and affordable small model for lightweight tasks and high-throughput use cases.
72
Quality
$0.38
Blended/1M
183
tok/s
Llama 4 Scout
Meta's efficient MoE model with 16 experts. 10M token context window and strong multilingual support.
71
Quality
$0.28
Blended/1M
198
tok/s
Amazon Nova Pro
Amazon's capable multimodal model. Strong balance of accuracy, speed, and cost for diverse tasks.
70
Quality
$2.00
Blended/1M
110
tok/s
Command R+
Cohere's flagship for enterprise RAG. Optimized for retrieval-augmented generation and tool use.
68
Quality
$6.25
Blended/1M
72
tok/s