AI Model Directory

Detailed specs, pricing, and benchmarks for every major AI model.

Anthropic
New

Claude Opus 5.5

Anthropic's 22 Sep 2026 flagship and the top model on Artificial Analysis Intelligence Index v4.3.2 at 58 (max effort, adaptive reasoning, with fallback) — five points clear of Claude Fable 5.1 and GPT-6 Astra at 53 on the same index, and first on the LMArena text board (1509, high effort, 25 Sep snapshot), though that interval is wide enough that it is not statistically separated from ranks two to four. Priced at $4/$20 per 1M tokens, 20% below Opus 5 and 60% below Fable 5.1 per token; Anthropic says it requires less compute to serve and that the price reflects it. Cache reads are $0.20, 0.05x the input rate, and the Batch API halves both rates. 1M context at standard pricing with no long-context surcharge, 128K maximum output (300K on the Batch API with a beta header). Thinking is adaptive and always on, with a default effort of medium. Cost per task depends heavily on the effort dial: Artificial Analysis measures $5.98 per Index task at max, $3.46 at xhigh, $1.82 at high and $1.34 at medium (scores 58/56/54/51) — so high effort already out-scores Fable 5.1 at max (53, $7.63) for under a quarter of the cost. The catch is verbosity: 260M output tokens across the Index against a 81M median, flagged 'very verbose'. Fast mode (research preview, Claude API only) is $8/$40. Anthropic's own benchmark table (vendor-run, with production safeguards on) has Terminal-Bench 4.0 at 66.4 against Fable 5.1's 55.8 and Astra's 57.9.

100

Quality

$12.00

Blended/1M

—

tok/s

Best for: Top-scoring frontier reasoning & agentic coding
Anthropic

Claude Fable 5

Anthropic's June 9 2026 flagship and the first generally available model in the Mythos class: the new #1 overall. State-of-the-art on nearly every tested benchmark: SWE-bench Pro 80.3% (vs Opus 4.8 69.2%, GPT-5.5 58.6%), FrontierCode Diamond 29.3%, OSWorld-Verified computer use 85.0%, and a leading GDPval-AA knowledge-work Elo of 1932. Routes high-risk cyber/bio prompts to Opus 4.8 as a safety fallback. $10/$50 pricing (2x Opus 4.8) with a 90% prompt-caching discount; exact context window and max output were not officially disclosed at launch.

99

Quality

$30.00

Blended/1M

58

tok/s

Best for: Frontier agentic coding & knowledge work
Anthropic

Claude Opus 5

Anthropic's 24 Jul 2026 flagship and, at launch, the highest-scoring verified model on the board (since superseded by Opus 5.5). A step change over Opus 4.8 on deep reasoning, long-horizon agentic work, and test-time compute scaling — at unchanged Opus pricing of $5/$25 per 1M tokens, roughly half the cost of Claude Fable 5. 1M context (default and maximum), 128K max output. Thinking is on by default rather than opt-in, the prompt-cache minimum drops to 512 tokens from 1024, and the full low-through-max effort ladder is available; low and medium effort are unusually strong here, which makes it cheaper to run than the sticker price suggests. Fast mode is Claude-API-only at $10/$50. Draws on a separate rate-limit pool from the Opus 4.x models.

99

Quality

$15.00

Blended/1M

74

tok/s

Best for: Frontier agentic coding & reasoning
Anthropic

Claude Mythos 5

Once the top-scoring model on the board, and the one almost nobody can call: Mythos 5 is available only through Anthropic's Project Glasswing, succeeding the invitation-only Mythos Preview. Capabilities, pricing, limits, and API behaviour are identical to Claude Fable 5 — $10/$50 per 1M tokens, 1M context, 128K output, always-on thinking with the raw chain of thought never returned, and a 30-day data-retention requirement that rules out zero-retention orgs. Listed here for completeness of the frontier; if you are not in Glasswing, Claude Opus 5 or Fable 5 is the model you can actually buy.

99

Quality

$30.00

Blended/1M

56

tok/s

Best for: Ceiling capability (limited access)
OpenAI

GPT-6 Astra

OpenAI's 3 Sep 2026 flagship and the first model in the GPT-6 generation. Joint #1 on the Artificial Analysis Intelligence Index v4.3 at launch at 53, tied with Claude Fable 5.1 (both since passed by Claude Opus 5.5 at 58 and Claude Sonnet 5.5 at 56 on v4.3.2) — but at roughly a third of Fable's cost per task ($3.26 vs $7.63 at max effort), because it reaches the same score with markedly fewer steps and fewer output tokens. Trained on OpenAI's largest run to date, over 100,000 GPUs at the Stargate site in Texas, and the first OpenAI model to use other models in a significant supervisory role during training. Set a record on DeepSWE v1.1 at 74%. Context is 1,050,000 tokens but the shape matters: 922K maximum input, 128K maximum output, and requests above 272K input tokens bill at 2x input and 1.5x output for the entire request, not just the overflow — a cliff that a million-token window actively invites you to walk off. Five reasoning-effort levels (low through max) act as a cost dial: 53/51/50/46 at max/high/medium/low against $3.26/$1.72/$1.54/$0.82 per task, which makes medium the sensible default for most workloads. Batch and Flex run at half the standard rate; Fast mode doubles it. Chat Completions, Responses and Batch only — no Realtime, Assistants, audio, embeddings or moderation. First OpenAI model to reach the company's internal 'Critical' cybersecurity threshold, scoring 100% on ExploitBench and finding two zero-days in a modified run, so the advanced cyber capabilities are access-gated via the Daybreak programme. Note for compliance teams: the model uses opaque recurrence, which OpenAI's chief scientist has said makes chain-of-thought auditing harder.

99

Quality

$30.00

Blended/1M

53

tok/s

Best for: Frontier software engineering & computer use
Anthropic

Claude Fable 5.1

Anthropic's 1 Sep 2026 refresh of Fable 5 and joint #1 on the Artificial Analysis Intelligence Index v4.3 at launch at 53, level with GPT-6 Astra (both since passed by Claude Opus 5.5 at 58 and Claude Sonnet 5.5 at 56 on v4.3.2). The two models also carry an identical sticker price of $10/$50, which makes everything else the deciding factor — and on cost per task Fable is the expensive one, at $7.63 against Astra's $3.26 for the same score, driven by verbosity (190M output tokens across the index versus a 90M median). The 98% cache discount is the counterweight: on workloads with heavy prompt reuse the blended rate falls to about $7.17 per 1M, and reuse-dominated traffic changes the arithmetic substantially. One hard constraint to design around: time to first token is 262.88s against a 3.73s median. That is thinking time rather than network latency, but it rules the model out of anything interactive regardless of the cause. 1M context, text and image in, text out, extended reasoning, served through 5 API providers. Anthropic has not published a parameter count.

99

Quality

$30.00

Blended/1M

68

tok/s

Best for: Top-tier reasoning with heavy prompt reuse
Anthropic
New

Claude Sonnet 5.5

Anthropic's 28 Sep 2026 mid-tier model, at $2/$10 per 1M tokens with cache reads at $0.20 (0.1x input). Artificial Analysis Intelligence Index v4.3.2 puts it at 56 at max effort, 52 at xhigh, 47 at high and 41 at medium — above Claude Fable 5.1 and GPT-6 Astra (both 53) at max, at a fifth of their per-token price. The cost-per-task figures are the counterweight: $7.60 per Index task at max (410M output tokens against an 81M median, flagged 'very verbose'), $2.74 at xhigh and $1.08 at high, so the effort setting matters more than the sticker price. Anthropic says it costs up to 30% less per task than Sonnet 5, runs 30%+ faster, and batches tool calls more; that comparison is not visible in Artificial Analysis's row for this model. Per-token pricing is unchanged from Sonnet 5 per Anthropic's pricing table. 1M context at standard pricing, 128K maximum output (300K on the Batch API with a beta header). Thinking is adaptive with a default effort of high and a new lowest setting, between_tools. The Batch API halves the rate to $1/$5. Fast mode is not offered.

99

Quality

$6.00

Blended/1M

—

tok/s

Best for: High-scoring agents when you can cap effort
Anthropic

Claude Opus 4.8

Anthropic's May 28 2026 flagship and the new #1 on the Artificial Analysis Intelligence Index at 61.4 (+4.1 over Opus 4.7, +1.2 ahead of GPT-5.5). SWE-bench Verified 88.6%, SWE-bench Pro 69.2%, Terminal-Bench 2.1 74.6%, and a leading 49.8% on Humanity's Last Exam. Strongest computer-use/browser-agent model tested (Online-Mind2Web 84%). Same $5/$25 pricing as Opus 4.7; fast mode is ~2.5x faster and ~3x cheaper.

98

Quality

$15.00

Blended/1M

72

tok/s

Best for: Coding, agents & computer use
OpenAI

GPT-5.6 Sol

OpenAI's flagship tier, GA 9 Jul 2026. The 5.6 generation replaced the old one-model-with-a-dial approach with three durable tiers — Sol (flagship), Terra (balanced), Luna (cheap) — and `gpt-5.6` is now an alias for `gpt-5.6-sol`. There is no separate Pro model ID; Pro is a reasoning mode on Sol. Artificial Analysis Intelligence Index 61.0, SWE-bench Verified 88.1%, SWE-bench Pro 66.4%, Terminal-Bench 2.1 72.9%. 1.05M context (922K input / 128K output). Pricing holds at $5/$30 with cache reads at 10% of input; requests over 272K tokens move to a long-context meter at $10/$45.

98

Quality

$17.50

Blended/1M

96

tok/s

Best for: General frontier reasoning & tools
OpenAI
New

GPT-6.1 Sol

OpenAI's 29 Sep 2026 refresh of GPT-6 Sol, at the same $2/$10 per 1M tokens but with cached input halved to $0.10. Artificial Analysis Intelligence Index v4.3.2 scores it 52 at max effort — one point under GPT-6 Astra (53) at $0.72 per Index task against Astra's $3.26 — and 51/50/48 at xhigh/high/medium for $0.39/$0.32/$0.21. OpenAI says it nearly matches Astra on agentic coding, computer use and professional work at one-fifth of Astra's standard token prices (vendor claim); OpenAI gave no reason for the price. It also uses far fewer tokens than the Claude 5.5 pair: 67M output tokens across the Index. 1,050,000 context with 922K maximum input and 128K maximum output; requests above 272K input tokens bill at 2x input and 1.5x output for the whole request. Effort levels low, medium (default), high, xhigh and max; text and image in, text out. Batch and Flex are half price and Fast mode doubles the rate. Knowledge cutoff 30 Apr 2026. An Ultrafast tier was announced for it but was not live on 29 Sep.

98

Quality

$6.00

Blended/1M

—

tok/s

Best for: Near-Astra agentic quality at a fifth of the price
OpenAI

GPT-5.5

OpenAI's "Spud": the first fully retrained base model since GPT-4.5. 1M context window; tops the Artificial Analysis Intelligence Index at 59-60 and the LMArena text leaderboard at ~1506 Elo. Shipped April 23 2026, on AWS Bedrock April 28.

97

Quality

$17.50

Blended/1M

70

tok/s

Best for: Frontier general purpose
Moonshot AI
OSS

Kimi K3

Moonshot AI's 16 Jul 2026 flagship and the clearest evidence that the open-weight tier has caught the frontier: a 2.8-trillion-parameter MoE (16 of 896 experts routed per token, Stable LatentMoE, MXFP4 weights) with a 1,048,576-token context and native text, vision, and video input. Artificial Analysis Intelligence Index 57, ranking #3 overall behind only Claude Fable 5 and GPT-5.6 Sol, and ahead of every other proprietary model on the board. It beats GLM-5.2 across Moonshot's harness — DeepSWE 67.5 vs 46.2, FrontierSWE 81.2 vs 67.3, SWE Marathon 42.0 vs 13.0, Terminal-Bench 2.1 88.3 vs 82.7, GPQA-Diamond 93.5 vs 91.2 — and tops the open board on Humanity's Last Exam (56%) and BrowseComp (91.2). At $3/$15 it undercuts both Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30) on input and output while scoring higher on independent benchmarks. Two caveats: K3 is a heavy token consumer, and it is a genuine cost outlier among open models at ~$0.94 per task versus DeepSeek V4 Pro's $0.04. Weights shipped on Hugging Face at moonshotai/Kimi-K3 (2.8T total, 104B active) under Moonshot's own Kimi K3 License rather than a standard open licence, so read the terms before building on it. Hallucination rate rose to 51% from K2.6's 39% — cross-check factual output.

97

Quality

$9.00

Blended/1M

55

tok/s

Best for: Open-weight coding frontier
Meta

Muse Spark 1.3

Meta's 2 Sep 2026 flagship and the clearest statement yet that the company has left open frontier releases behind: Muse Spark 1.3 is closed, Meta has published no parameter count, and Llama 5 has slipped to roughly 2027. Taken on its own terms it is a strong model and a genuinely fast one. Artificial Analysis rates it 48 on the v4.3 Intelligence Index — above GPT-5.6 Sol at 47 and behind only the Astra, Fable 5.1 and Opus 5 families — at $1.25/$4.25, which is an eighth of Astra's input price. Output speed is 255.2 tokens/second, fifth of 199 models measured and nearly four times the 68 tok/s median. The offsetting figure is a 27.01s time to first token, so it is fast once it starts but not quick to start; for batch and throughput-bound work that trade is usually worth taking, for interactive work it usually is not. 1M context. Speed and latency figures are from Meta's first-party API.

97

Quality

$2.75

Blended/1M

255

tok/s

Best for: High-throughput batch inference
OpenAI
New

GPT-6 Sol

OpenAI's 22 Sep 2026 mid-tier GPT-6 model at $2/$10 per 1M tokens, cached input $0.20. OpenAI's saving was a rate cut rather than fewer tokens: reported as exactly 50% below GPT-5.6 Sol in both directions, with Artificial Analysis noting slightly more output tokens per task (77M across the Index). Artificial Analysis Intelligence Index v4.3.2: 48 at max and 44 at xhigh, at $1.05 and $0.52 per Index task. Superseded a week later by GPT-6.1 Sol at the same price and a higher score (52 at max). 1,050,000 context with 922K maximum input and 128K maximum output per OpenAI's docs (Artificial Analysis lists 872k); requests above 272K input tokens bill at 2x input and 1.5x output for the whole request. Knowledge cutoff 20 Apr 2026. Batch and Flex are half price and Fast mode doubles the rate.

97

Quality

$6.00

Blended/1M

—

tok/s

Best for: Low-cost GPT-6 generation (see GPT-6.1 Sol)
OpenAI

GPT-5.5 Pro

High-compute variant of GPT-5.5 with extended thinking. 6x the price for mid-teens AAII uplift on hard reasoning workloads.

96

Quality

$105.00

Blended/1M

68

tok/s

Best for: Reasoning at any cost
Anthropic

Claude Opus 4.7

Anthropic's April 16 2026 flagship. SWE-bench Verified 87.6%, SWE-bench Pro 64.3%, GPQA Diamond 94.2%. Sits in the LMArena top tier at ~1503 (1505 thinking). Same $5/$25 pricing as Opus 4.6.

96

Quality

$15.00

Blended/1M

68

tok/s

Best for: Coding & agentic workflows
Google

Gemini 3.1 Pro

Tied at the top of the LMArena text leaderboard May 2026 @ ~1505 Elo. GPQA Diamond 94.3%, MMLU-Pro 91.0%: current scientific reasoning leader. 1M token context; fastest frontier model at ~131 tok/s.

96

Quality

$7.00

Blended/1M

131

tok/s

Best for: Science & long-context
Alibaba Cloud
OSS

Qwen3.8 Max

Reached GA on 3 Aug 2026 and, around 12–14 Aug, became the first Qwen-Max-class model Alibaba has ever open-weighted — following its 19 Jul preview at WAIC Shanghai. Still a 2.4T-parameter sparse MoE (active-parameter count undisclosed), but treat the hosted API and the open weights as two different products: the published checkpoint is reportedly text-only, dropping the vision and 1M-context features the hosted API keeps. Weights ship under a new custom revenue-share license, not Apache 2.0 — large commercial 'model-as-a-service' deployments above an unspecified threshold need a separate agreement with Alibaba, so check the license before redistributing. Vendor-reported scores at GA: Terminal-Bench 2.1 86.6 (vs. Claude Opus 4.8/Fable 5 at 84.6, GPT-5.6 Sol at 88.8), SWE-bench Pro 67.7, GPQA Diamond 92.6, DeepSWE 1.1 56.6 (up from 21.6 on Qwen3.7 Max) — independent Arena/AA verification is still catching up as of 17 Aug 2026, though early Arena placement (~1491) already makes it the highest-ranked open-weight model on the board. Hosted API pricing dropped from the preview placeholder to $2/$6 per 1M tokens ($0.25 cached).

96

Quality

$4.00

Blended/1M

86

tok/s

Best for: Multimodal APAC frontier (hosted); text-only open weights
Z.ai (Zhipu AI)
OSS

GLM-5.3

Z.ai's 18 Aug 2026 flagship, held back from the 17 Aug catalogue review because the weights had been announced on 14 Aug but not yet published. They are now on Hugging Face: a 753B-parameter MoE with 40B active per token, under Z.ai's own GLM-5.3 License rather than a standard open licence — commercial use is permitted subject to restrictions, so read the terms before redistributing. Artificial Analysis rates it 45 on the v4.3 Intelligence Index, the highest of any comparable open-weight model, against a class median of 18. The catch is price: at $1.40/$4.40 it is, in Artificial Analysis's own words, 'particularly expensive when comparing to other open weight models of similar size' — and its own stablemate GLM-5.3-Flash scores 42 for a tenth of the input cost. Speed is unremarkable at 67.0 tokens/second, slightly below the 69.3 class median, with a 3.11s time to first token. 1M context, text in and text out only — no vision, unlike the Flash variant. Served by 22 API providers, the widest availability of any open-weight model here.

96

Quality

$2.90

Blended/1M

67

tok/s

Best for: Highest-scoring open-weight flagship
xAI

Grok 4.6

Released 12 Aug 2026, scoring 44 on the Artificial Analysis Intelligence Index v4.3 at high and xhigh effort (43 at medium), rank #21 of 199, at $2/$6. The notable spec is the one that moved the wrong way relative to the field: a 500k context window, against 1M for almost every other frontier model on this board. For most work that is academic — 500k is around 750 A4 pages — but it rules out the whole-repository and whole-corpus patterns that the 1M-context models have made routine. Output speed is 64.5 tokens/second, just under the 68 tok/s median, with a 75% cache discount bringing the blended rate to about $1.35 per 1M. Text and image in, text out, reasoning model. Note that Artificial Analysis now lists the creator as SpaceXAI.

95

Quality

$4.00

Blended/1M

65

tok/s

Best for: Balanced reasoning under 500k context
OpenAI

o3

OpenAI's most powerful reasoning model. State-of-the-art on MATH, coding, and science benchmarks.

94

Quality

$5.00

Blended/1M

68

tok/s

Best for: Hard reasoning
Alibaba Cloud

Qwen 3.7 Max

Alibaba's May 20 2026 proprietary flagship, unveiled at the Alibaba Cloud Summit in Hangzhou. Highest-ranked Chinese model on the Artificial Analysis Intelligence Index at 56.6 (#5 overall, +4.8 over Qwen 3.6 Max Preview). 1M context; SWE-bench Pro 60.6, Terminal-Bench 2.0 69.7, GPQA Diamond 92.4, and a table-leading 97.1 on HMMT Feb 2026 competition math. Ran 35 hours autonomously across 1,158 tool calls and supports external harnesses like Claude Code.

94

Quality

$5.00

Blended/1M

90

tok/s

Best for: Long autonomous agentic runs
xAI

Grok 4.5

xAI's July 2026 flagship at $2 input / $6 output per 1M tokens — notable for keeping a 3:1 output-to-input ratio where most rivals charge 5-6x, which makes it materially cheaper than headline comparisons suggest on generation-heavy work. 500K context, Artificial Analysis Intelligence Index 54. Watch three things: prompts over 200K tokens double in price, Grok 4.5 is excluded from the batch discount that applies to Grok 4.3, and tools bill separately (Web/X Search and Code Execution at $5 per 1,000 calls, File Attachments at $10 per 1,000). Cached input is quoted between $0.30 and $0.50 depending on source — verify in the console before modelling costs.

94

Quality

$4.00

Blended/1M

88

tok/s

Best for: Cheap output tokens & real-time info
xAI

Grok 4.3

xAI's May 6 2026 flagship. 1M context, native video input, and a ~40% input price cut over Grok 4.20. Artificial Analysis Intelligence Index 53; outperforms Opus 4.7 ~1.26x on agentic Vending-Bench.

93

Quality

$1.88

Blended/1M

83

tok/s

Best for: Agentic tasks & real-time info
OpenAI

GPT-5.6 Terra

The balanced everyday tier of the GPT-5.6 family, GA 9 Jul 2026. On 30 Jul 2026 OpenAI cut Terra 20% to $2 input / $12 output per 1M tokens, crediting inference work that reduced end-to-end serving cost by 20% and improved token-generation efficiency by more than 15%. Same 1.05M context as Sol, cache reads at $0.20, Batch API at half rate. This is the tier most production workloads should default to — it lands within a few points of Sol on most evals at 40% of the price.

93

Quality

$7.00

Blended/1M

118

tok/s

Best for: Default production tier
Anthropic

Claude Sonnet 5

Anthropic's 25 Jun 2026 mid-tier model and the workhorse of the Claude 5 family. It carries most of the Opus-line coding and agent quality at roughly a third of the price: SWE-bench Verified 82.4%, Terminal-Bench 2.1 68.1%, Artificial Analysis Intelligence Index 57.2. List price is $3/$15 with a 90% caching discount, but an introductory rate of $2/$10 runs through 31 Aug 2026 — worth locking in evaluation work before it lapses. 1M-token context and the same tool-use API as the flagship line. The default pick for high-volume agent work where Opus 5 or Fable 5 would be overkill on cost.

93

Quality

$9.00

Blended/1M

98

tok/s

Best for: Balanced agents & coding value
Z.ai (Zhipu AI)
OSS

GLM-5.3 Flash

Released 26 Aug 2026 under a straight MIT licence and, on price per unit of measured intelligence, the strongest value on the board. A 320B-parameter MoE with 18B active, scoring 42 on the Artificial Analysis Intelligence Index v4.3 — three points behind its own 753B flagship GLM-5.3, at $0.15/$0.50 against that model's $1.40/$4.40. That is roughly a tenth of the input cost for 93% of the score. It is also the open-weight model most teams can realistically self-host: 320B total is demanding but not rack-scale in the way Kimi K3's 2.8T is, and the 18B active count keeps inference cheap. 1M context, text and image in, text out, 117.2 tokens/second, 18 API providers. One caveat: Artificial Analysis flags it 'very verbose' at 180M output tokens across the index against a 140M median, so budget output tokens at roughly 1.3x what the headline rate implies.

93

Quality

$0.33

Blended/1M

117

tok/s

Best for: Best open-weight price-to-performance
Google

Gemini 3.8 Flash

Released 2 Sep 2026 alongside a restricted Gemini 3.8 Flash Cyber variant — by Google's own count the third Flash release in six weeks, and further evidence that the Pro line has stalled: Gemini 3.1 Pro (19 Feb 2026) is still the current Pro model, and every release since has been Flash-tier. The headline number is speed. At 345.1 tokens per second it is the fastest model of the 199 Artificial Analysis measures, roughly five times the 68 tok/s median, while scoring 41 on the v4.3 Intelligence Index. The offset is a 30.49s time to first token, so the throughput only pays off on long generations. Pricing holds at the 3.7 Flash introductory rate of $0.75/$3.75 with a 90% cache discount — but note that this is introductory: from 1 Jan 2027 the rates double to $1.50 and $7.50, which is worth modelling now if you are sizing a 2027 budget. Google's own benchmark table has 3.8 Flash ahead of 3.7 Flash on every published row: DeepSWE v1.1 73.7% vs 65.3%, Terminal-bench 2.1 89.4% vs 85.8%, OSWorld-2.0 59.0% vs 50.6%, Vals Finance Agent v2 61.4% vs 59.0%, HLE-Verified 54.9% vs 53.6% — all first-party runs. Google attributes the gains to the model executing extra reasoning steps and calling tools iteratively, which also explains why Artificial Analysis flags it 'very verbose' at 170M output tokens against a 90M median. 1M context. The Cyber variant is gated behind the new Fairwind Programme for government authorities, critical infrastructure operators and software maintainers.

93

Quality

$2.25

Blended/1M

345

tok/s

Best for: Fastest high-throughput generation
Google

Gemini 2.5 Pro

Google's thinking model with native tool use, 1M context window, and strong multimodal capabilities.

92

Quality

$5.63

Blended/1M

87

tok/s

Best for: Multimodal + value
Moonshot AI

Kimi K2.6

Moonshot AI's April 20 2026 frontier model. 256K context with text, image, and video input. Artificial Analysis Intelligence Index 54; SWE-bench Verified 80.2%, up sharply from K2.5.

92

Quality

$2.11

Blended/1M

48

tok/s

Best for: Frontier quality at low cost
DeepSeek
NewOSS

DeepSeek V4.1 Flash

Released 10 Sep 2026 with weights on Hugging Face under a genuine MIT licence — the permissive one, with no revenue-share clause, no territorial exclusion and no custom terms to read. That matters because it posts Terminal-Bench 2.1 of 90.6, ahead of Claude Opus 5 (89.1) and GPT-5.6 Sol (88.8), at $0.30/$1.20 hosted. The smallest model in DeepSeek's new architecture family: a multimodal MoE with a 552B backbone but only 8B parameters active while processing a prompt and 16B while generating, using a Causal Encoder-Decoder design — 40 layers arranged as a 20-layer causal encoder followed by a 20-layer decoder, where the decoder's global KV cache is projected from the final encoder hidden states rather than from each decoder layer's own. KV entries are stored in four-bit floating point at roughly 890 bytes per token, about a quarter of what V4 Flash needed, and the persistent SSD cache drops to around an eighth of the previous generation. 1M context, native image and text in. Two caveats worth pricing in: Artificial Analysis rates it 40 on the v4.3 Intelligence Index — below GLM-5.3 (45) and Kimi K3 (44) on the composite even as it leads them on Terminal-Bench — and it is flagged 'very verbose' at roughly 250M output tokens across the index against a 140M median, which quietly multiplies that cheap output rate. Call it as `deepseek-flash`; V4 Flash and V4 Flash Vision Exp are retired and their old names route here temporarily. DeepSeek had planned to route all `deepseek-v4-pro` traffic here from 14 Sep but reversed that after user pushback, so V4 Pro continues with billing unchanged. Available through 8 API providers.

92

Quality

$0.75

Blended/1M

217

tok/s

Best for: Cheap open-weight agentic coding
Anthropic

Claude Opus 4

Anthropic's most capable model. Excels at complex analysis, nuanced writing, and extended agentic tasks.

91

Quality

$45.00

Blended/1M

52

tok/s

Best for: Complex analysis
DeepSeek
OSS

DeepSeek R1

DeepSeek's reasoning model. Competitive with o3 on math and coding at a fraction of the cost.

91

Quality

$1.37

Blended/1M

35

tok/s

Best for: Cheap reasoning
Z.ai (Zhipu AI)
OSS

GLM-5.2

Z.ai's 13 Jun 2026 open-weight flagship, and the model that held the top open-weight slot until Kimi K3 landed a month later. 744B MoE / ~40B active, MIT-licensed weights on Hugging Face, 1M context, 131K max output. $1.40 input / $4.40 output per 1M tokens ($0.26 cached) — a blended ~$0.90/1M. The headline result is SWE-bench Pro 62.1%, which beats GPT-5.5's 58.6%: an MIT-licensed model you can self-host outscoring a $5/$30 US flagship on agentic coding. Also the fastest of the big open MoEs at ~168 tokens/sec, roughly 2.7x Kimi K3 and DeepSeek V4 Pro. Self-hosting needs about 1TB VRAM in BF16, or ~8x H200 at FP8.

91

Quality

$2.90

Blended/1M

168

tok/s

Best for: Fastest self-hostable frontier model
Anthropic

Claude Sonnet 4.6

Anthropic's Feb 17 2026 balanced model. Near-Opus performance at Sonnet pricing: SWE-bench Verified 79.6%, OSWorld 72.5%, 1M context. Best price-to-performance in the Claude lineup.

90

Quality

$9.00

Blended/1M

73

tok/s

Best for: Coding & balance
Google

Gemini 3.6 Flash

Google's 21 Jul 2026 mid-tier release, replacing Gemini 3.5 Flash. Introductory pricing of $0.75 input / $3.75 output per 1M tokens runs through 31 Dec 2026, after which the standard $1.50 / $7.50 rate applies; cached input is $0.15 and batch mode is half rate. 1M context, native audio and video understanding, and the fastest response times in the frontier-adjacent band. Note that Google shipped no Gemini 3.2–3.5 Pro: the Pro line still tops out at Gemini 3.1 Pro, and the 3.5/3.6 releases are all Flash-tier.

90

Quality

$2.25

Blended/1M

148

tok/s

Best for: Fast multimodal at mid-tier price
OpenAI

GPT-4.1

OpenAI's latest flagship with 1M token context, improved instruction following and coding.

89

Quality

$5.00

Blended/1M

120

tok/s

Best for: Long context
DeepSeek
OSS

DeepSeek V4 Pro

Left preview and reached GA on 12–13 Aug 2026 as the '0813' checkpoint — same 1.6T MoE / 49B-active architecture as the April launch, but re-post-trained and, notably, repriced upward. The promotional $0.435/$0.87 rate is gone: GA pricing is $0.66 input / $1.98 output per 1M tokens, a roughly 52% input and 128% output increase, which DeepSeek framed as stepping back from a pure race-to-zero strategy. Still MIT-licensed weights on Hugging Face, 1M context. Vendor-reported scores are strong — LiveCodeBench 93.5%, SWE-bench Verified 80.6% (matching Gemini 3.1 Pro), GPQA Diamond 90.1%, Codeforces 3206 — but treat them cautiously: contamination-resistant independent benchmarks (DeepSWE) and Code Arena both show a wider gap to closed frontier models than DeepSeek's own numbers suggest, and third-party Arena Elo for the 0813 build is still preliminary (~1458–1465) as of 17 Aug 2026.

89

Quality

$1.32

Blended/1M

62

tok/s

Best for: Cheap frontier-adjacent tokens (GA build)
MiniMax
OSS

MiniMax M3

MiniMax's 1 Jun 2026 open-weight flagship: the first open model to combine frontier agentic coding, native multimodality, and a 1M-token context. Built on MiniMax Sparse Attention (MSA): ~1/20th the per-token compute at 1M context, >9x faster prefill, >15x faster decode. Now third-party verified and holding up well — SWE-bench 80.5% (second among open models, behind only DeepSeek V4 Pro's 80.6%), GPQA Diamond 93.0% (second behind Kimi K3), OSWorld 70.1%, BrowseComp 83.5%, Terminal-Bench 2.1 66.0%. At $0.60/$2.40 it delivers SWE-bench parity with models costing 10x more. Launch-promo pricing was $0.30/$1.20.

89

Quality

$1.50

Blended/1M

80

tok/s

Best for: Open-weight agentic coding
OpenAI
New

GPT-6 Luna

OpenAI's 22 Sep 2026 small GPT-6 model at $0.10/$0.50 per 1M tokens with cached input at $0.01. Artificial Analysis Intelligence Index v4.3.2 scores it 37 at max effort at $0.07 per Index task. It is verbose for its size: 140M output tokens across the Index, so the low rate is partly spent on length. 1,050,000 context with 922K maximum input and 128K maximum output; requests above 272K input tokens bill at 2x input and 1.5x output for the whole request. Knowledge cutoff 18 May 2026. Batch and Flex are half price and Fast mode doubles the rate.

89

Quality

$0.30

Blended/1M

—

tok/s

Best for: Cheapest current GPT-6 for high-volume work
OpenAI

o3 Mini

OpenAI's compact reasoning model with extended thinking capabilities for complex problem solving.

88

Quality

$2.75

Blended/1M

155

tok/s

Best for: Reasoning & math
Anthropic

Claude Sonnet 4

Anthropic's balanced model with excellent coding and reasoning. Best price-to-performance ratio.

88

Quality

$9.00

Blended/1M

95

tok/s

Best for: Coding & balance
Z.ai (Zhipu AI)
OSS

GLM-5.1

Z.ai's (Zhipu) 2026 flagship. 200K context, strong agentic and tool-use scores: 98% on τ²-Bench Telecom, on par with Grok 4.3. Artificial Analysis Intelligence Index 51.

88

Quality

$2.03

Blended/1M

48

tok/s

Best for: Open-weight agentic & tool use
xAI

Grok 3

xAI's flagship model with strong reasoning and real-time information access. Trained on the Colossus cluster.

87

Quality

$9.00

Blended/1M

82

tok/s

Best for: Real-time info
DeepSeek
OSS

DeepSeek V3

671B MoE model with 37B active parameters. Outstanding price-performance ratio and coding ability.

86

Quality

$0.69

Blended/1M

62

tok/s

Best for: Best open-source value
Alibaba Cloud

Qwen 3.6 Plus

Alibaba's April 2026 flagship. Long-context multilingual specialist with strong CN/EN/AR coverage and competitive coding scores.

86

Quality

$3.50

Blended/1M

124

tok/s

Best for: Multilingual & APAC
OpenAI

GPT-4o

OpenAI's flagship multimodal model with vision, code generation, and function calling. Excellent all-round performance.

85

Quality

$6.25

Blended/1M

109

tok/s

Best for: General purpose
DeepSeek
OSS

DeepSeek V4 Flash

The cheapest usable model in the directory at $0.14 input / $0.28 output per 1M tokens, with cache hits at $0.0028 — DeepSeek cut the cache-hit rate to a tenth of its launch price on 26 Apr 2026. 284B MoE / 13B active, MIT weights, 1M context, and an OpenAI-compatible endpoint. On 31 Jul 2026 the `deepseek-v4-flash` API ID began serving the V4-Flash-0731 checkpoint, which posts 79% on SWE-bench Verified and 85.9 on BrowseComp — numbers that would have been frontier a year earlier, at roughly 1/100th of GPT-5.6 Sol's output price. The hosted service is in public beta. Rate limit is 2,500 concurrent requests.

84

Quality

$0.21

Blended/1M

105

tok/s

Best for: Cheapest tokens, period
OpenAI

GPT-5.6 Luna

The fastest, cheapest member of the GPT-5.6 family and the site of the steepest price cut of the year from a US lab: on 30 Jul 2026 OpenAI dropped Luna 80%, from $1/$6 to $0.20 input / $1.20 output per 1M tokens. Cache reads run $0.02. That is an explicit answer to DeepSeek and the Chinese open-weight tier on price — though at $0.20/$1.20 Luna is still comfortably more expensive than DeepSeek V4 Flash at $0.14/$0.28. Full 1.05M context, unlike most rivals' cheap tiers.

83

Quality

$0.70

Blended/1M

186

tok/s

Best for: High-volume cheap frontier-family
Meta
OSS

Llama 4 Maverick

Meta's mixture-of-experts model with 17B active parameters and 128 experts. Strong multimodal and multilingual performance.

80

Quality

$0.40

Blended/1M

135

tok/s

Best for: Open-source value
Alibaba Cloud
OSS

Qwen 2.5 72B

Alibaba's flagship open-source model. Competitive with GPT-4o class models on benchmarks at a fraction of the cost.

80

Quality

$0.60

Blended/1M

85

tok/s

Best for: Open-source flagship
Alibaba Cloud
OSS

Qwen3.8 27B

Alibaba's fully open, Apache 2.0-licensed sibling to Qwen3.8 Max, released 13–14 Aug 2026 — the locally-runnable pick where the Max tier's custom revenue-share license doesn't apply. 27.8B dense params (not MoE) with native vision-language input and Hybrid Gated DeltaNet attention (3-of-4 sublayers linear attention), giving a native 262K context extensible to 1M via YaRN RoPE scaling. Vendor-reported agentic-coding gains over its predecessor are large: SWE-bench Pro 61.7, DeepSWE 1.1 42.2 (vs. 13.3 on Qwen3.6 27B), Terminal-Bench 2.1 73.0, LiveCodeBench v6 90.3, OSWorld 84.3 (vs. 63.9) — Alibaba claims it beats Claude Opus 4.6 Max on several of these. Runs in roughly 17GB VRAM at 4-bit, the main reason to reach for it over Max for self-hosting; one independent review found it markedly slower and more token-hungry per task than its predecessor at inference, so budget for that trade-off. No independent MMLU/Arena placement yet as of 17 Aug 2026.

80

Quality

$0.00

Blended/1M

—

tok/s

Best for: Locally-runnable open-weight (~17GB VRAM at 4-bit)
Mistral AI

Mistral Large 2

Mistral's flagship 123B model with strong multilingual and coding performance. Supports 128K context.

79

Quality

$4.00

Blended/1M

78

tok/s

Best for: Multilingual
Google

Gemini 3.5 Flash-Lite

Google's cheapest current-generation model at $0.30 input / $2.50 output per 1M tokens ($0.03 cached), released 21 Jul 2026 alongside 3.6 Flash. 1M context. Competitive on high-volume classification and extraction work, though DeepSeek V4 Flash still undercuts it roughly 2x on input and 9x on output.

79

Quality

$1.40

Blended/1M

192

tok/s

Best for: High-volume cheap multimodal
NVIDIA
OSS

Nemotron 3.5 Lightning

NVIDIA's 11 Aug 2026 release, the smallest model in the Nemotron 3 family (sibling to Nemotron 3 Ultra, distinct from the catalog's Nemotron 3 Nano Omni). Hybrid Mamba-2 + MoE + attention architecture, 30B total / 3B active params, pretrained on 20T+ tokens with an NVFP4 recipe, under NVIDIA's permissive OpenMDW-1.1 license (open weights, training data, and recipes; commercial use allowed). 1M-token context. Ships with multi-token-prediction draft models (DSpark, DFlash) for speculative decoding and is single-GPU deployable (1x DGX Spark or 1x H100). Alongside it NVIDIA released NeMo Switchyard, an open-source agent-routing library. Vendor-reported: MMLU Pro 81.9, GPQA Diamond 75.4, SWE-bench Verified 51.6-52.8, PinchBench 86% accuracy (30% faster than Qwen3.6 35B). No independent Arena placement yet as of 17 Aug 2026; hosted via build.nvidia.com and OpenRouter, but paid-tier pricing wasn't confirmed at the time of writing — verify before quoting a rate.

79

Quality

$0.00

Blended/1M

—

tok/s

Best for: Efficient single-GPU deployment, 1M context
xAI

Grok 3 Mini

xAI's efficient reasoning model with thinking capabilities at a lower cost point.

78

Quality

$0.40

Blended/1M

165

tok/s

Best for: Budget reasoning
Perplexity

Sonar Pro

Perplexity's search-augmented model. Combines LLM reasoning with real-time web search and citations.

78

Quality

$9.00

Blended/1M

65

tok/s

Best for: Search + citations
Mistral AI

Codestral

Mistral's dedicated code model. Optimized for code generation, completion, and review across 80+ languages.

76

Quality

$0.60

Blended/1M

195

tok/s

Best for: Code generation
NVIDIA
OSS

Nemotron 3 Nano Omni

NVIDIA's 30B open multimodal model running vision + audio + text in a single stack. Tops 6 specialty leaderboards. April 2026.

76

Quality

$0.00

Blended/1M

158

tok/s

Best for: Open multimodal
Tencent
OSS

Hunyuan Hy3

Tencent's 6 Jul 2026 full release (previewed 23 Apr 2026), a 295B-total / 21B-active MoE plus a separate 3.8B multi-token-prediction layer, under Apache 2.0 with no field-of-use or geographic restrictions. 256K native context. Tencent concedes the agentic-coding crown to GLM-5.2 — SWE-bench Pro rose from 46.0 at preview to 57.9 at release, still behind GLM — but claims leads in agentic search, tool use, and long-context retrieval at under half GLM-5.2's memory footprint. First Tencent entry in this catalog; Hunyuan's flagship chat line (outside this Hy3 open release) remains closed-weight/API-only. No independent Arena placement yet as of 17 Aug 2026.

76

Quality

$0.00

Blended/1M

—

tok/s

Best for: Long-context retrieval & agentic search, no license restrictions
Anthropic

Claude 3.5 Haiku

Anthropic's fastest model. Ultra-low latency for real-time applications and high-volume tasks.

75

Quality

$2.40

Blended/1M

172

tok/s

Best for: Speed & cost
Google
OSS

Gemma 4 27B

Google's open-weight flagship under Apache 2.0: April 2026. Designed for self-host with strong instruction-following and tool calling.

75

Quality

$0.00

Blended/1M

142

tok/s

Best for: Self-hosted general purpose
Google

Gemini 2.0 Flash

Google's fastest model. Optimized for speed and efficiency with strong coding and reasoning.

74

Quality

$0.25

Blended/1M

244

tok/s

Best for: Fastest + cheapest
Alibaba Cloud
OSS

Qwen 2.5 Coder 32B

Specialized coding model from Alibaba. Top open-source code model on HumanEval and SWE-Bench.

74

Quality

$0.30

Blended/1M

125

tok/s

Best for: Open-source coding
Ant Group
OSS

Ling-3.0-Flash

First catalog entry for Ant Group's inclusionAI/Bailing team, an actively-shipping open-weight lab (Ling non-thinking, Ring thinking, Ming multimodal families) not previously tracked here despite multiple trillion-parameter MIT-licensed releases through 2026 (Ling-2.5-1T, Ring-2.5-1T, Ling-2.6-1T). Ling-3.0-Flash, announced 27 Jul 2026 with weights open-sourced 5 Aug 2026, is a 124B-total MoE with roughly 5-7B active params per token (sources give both figures; verify against the official Hugging Face card before treating either as final), 256K context, MIT license. Ant claims it matches or beats their own 1T-parameter Ling-2.6-1T flagship on most benchmarks at roughly 1/8th the total parameters — a notable efficiency claim, but independent verification and a standard benchmark table weren't available at the time of writing. Weights ship in BF16 (255GB) and FP8 (128GB) on Hugging Face and ModelScope.

74

Quality

$0.00

Blended/1M

—

tok/s

Best for: Efficient MoE alternative to trillion-param flagships
OpenAI

GPT-4o Mini

Fast and affordable small model for lightweight tasks and high-throughput use cases.

72

Quality

$0.38

Blended/1M

183

tok/s

Best for: High throughput
Meta
OSS

Llama 4 Scout

Meta's efficient MoE model with 16 experts. 10M token context window and strong multilingual support.

71

Quality

$0.28

Blended/1M

198

tok/s

Best for: Longest context
Amazon

Amazon Nova Pro

Amazon's capable multimodal model. Strong balance of accuracy, speed, and cost for diverse tasks.

70

Quality

$2.00

Blended/1M

110

tok/s

Best for: AWS ecosystem
Cohere

Command R+

Cohere's flagship for enterprise RAG. Optimized for retrieval-augmented generation and tool use.

68

Quality

$6.25

Blended/1M

72

tok/s

Best for: Enterprise RAG

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.