Updated Oct 4, 2026

AI Model Leaderboard, October 2026

Every major AI model ranked by quality, speed, pricing, and value. Filter by category, sort by any metric, and find the right model for your use case. Live data refreshed regularly with LMSys Arena Elo, official provider pricing, and Artificial Analysis benchmarks.

In short:As of October 2026, Claude Opus 5.5 leads the AI model leaderboard at 100/100 on our composite quality index, across 69+ models ranked by quality, price, speed, and value. The best pick depends on the workload: sort the table by the metric that matters to you.

🥇

Claude Opus 5.5

100/100

🥈

Claude Fable 5

99/100

🥉

Claude Opus 5

99/100

Climb the Leaderboard

Stop reading: start ranking

Monthly Snapshot

October 2026: Top Models, Best Value, Fastest Inference

The October 2026 ranking covers 69 models across LMSys Arena Elo, MMLU Pro, HumanEval, MATH, pricing, and inference speed. Top of the table: Claude Opus 5.5 at 100/100 quality. The full table below is sortable by any metric. Live data is refreshed regularly from official provider pricing pages and the public Arena.

Top 5 by Quality Index

  1. Claude Opus 5.5: 100/100
  2. Claude Fable 5: 99/100
  3. Claude Opus 5: 99/100
  4. Claude Mythos 5: 99/100
  5. GPT-6 Astra: 99/100

Best Price-to-Quality

  1. DeepSeek V4 Flash: $0.28/1M out
  2. GLM-5.3 Flash: $0.5/1M out
  3. Gemini 2.0 Flash: $0.4/1M out
  4. GPT-6 Luna: $0.5/1M out
  5. Llama 4 Scout: $0.4/1M out

See our LMSys Arena deep dive and the monthly release roundup.

69 models
#ModelQualityArena ELOSpeedPriceContextValueReleased
1

Anthropic · Top-scoring frontier reasoning & agentic coding

100
1509—$4 / $201M8.3Sep 2026
2

Anthropic · Frontier agentic coding & knowledge work

99
152558 t/s$10 / $501M3.3Jun 2026
3

Anthropic · Frontier agentic coding & reasoning

99
152274 t/s$5 / $251M6.6Jul 2026
4

Anthropic · Ceiling capability (limited access)

99
153156 t/s$10 / $501M3.3Jun 2026
5

OpenAI · Frontier software engineering & computer use

99
—53 t/s$10 / $501M3.3Sep 2026
6

Anthropic · Top-tier reasoning with heavy prompt reuse

99
—68 t/s$10 / $501M3.3Sep 2026
7

Anthropic · High-scoring agents when you can cap effort

99
——$2 / $101M16.5Sep 2026
8

Anthropic · Coding, agents & computer use

98
151272 t/s$5 / $251M6.5May 2026
9

OpenAI · General frontier reasoning & tools

98
151496 t/s$5 / $301M5.6Jul 2026
10

OpenAI · Near-Astra agentic quality at a fifth of the price

98
——$2 / $101M16.3Sep 2026
11

OpenAI · Frontier general purpose

97
150670 t/s$5 / $301M5.5Apr 2026
12

Moonshot AI · Open-weight coding frontier

97
150055 t/s$3 / $151M10.8Jul 2026
13

Meta · High-throughput batch inference

97
—255 t/s$1.25 / $4.251M35.3Sep 2026
14

OpenAI · Low-cost GPT-6 generation (see GPT-6.1 Sol)

97
——$2 / $101M16.2Sep 2026
15

OpenAI · Reasoning at any cost

96
151068 t/s$30 / $1801M0.9Apr 2026
16

Anthropic · Coding & agentic workflows

96
150568 t/s$5 / $251M6.4Apr 2026
17

Google · Science & long-context

96
1505131 t/s$2 / $121M13.7Apr 2026
18

Alibaba Cloud · Multimodal APAC frontier (hosted); text-only open weights

96
149186 t/s$2 / $61M24.0Aug 2026
19

Z.ai (Zhipu AI) · Highest-scoring open-weight flagship

96
—67 t/s$1.4 / $4.41M33.1Aug 2026
20

xAI · Balanced reasoning under 500k context

95
—65 t/s$2 / $6500K23.8Aug 2026
21

OpenAI · Hard reasoning

94
137068 t/s$2 / $8200K18.8Apr 2025
22

Alibaba Cloud · Long autonomous agentic runs

94
148890 t/s$2.5 / $7.51M18.8May 2026
23

xAI · Cheap output tokens & real-time info

94
149988 t/s$2 / $6500K23.5Jul 2026
24

xAI · Agentic tasks & real-time info

93
149683 t/s$1.25 / $2.51M49.6May 2026
25

OpenAI · Default production tier

93
1489118 t/s$2 / $121M13.3Jul 2026
26

Anthropic · Balanced agents & coding value

93
147998 t/s$3 / $151M10.3Jun 2026
27

Z.ai (Zhipu AI) · Best open-weight price-to-performance

93
—117 t/s$0.15 / $0.51M286.2Aug 2026
28

Google · Fastest high-throughput generation

93
—345 t/s$0.75 / $3.751M41.3Sep 2026
29

Google · Multimodal + value

92
134587 t/s$1.25 / $101M16.4Mar 2025
30

Moonshot AI · Frontier quality at low cost

92
146648 t/s$0.73 / $3.49256K43.6Apr 2026
31

DeepSeek · Cheap open-weight agentic coding

92
—217 t/s$0.3 / $1.21M122.7Sep 2026
32

Anthropic · Complex analysis

91
136052 t/s$15 / $75200K2.0May 2025
33

DeepSeek · Cheap reasoning

91
135035 t/s$0.55 / $2.19128K66.4Jan 2025
34

Z.ai (Zhipu AI) · Fastest self-hostable frontier model

91
1483168 t/s$1.4 / $4.41M31.4Jun 2026
35

Anthropic · Coding & balance

90
146773 t/s$3 / $151M10.0Feb 2026
36

Google · Fast multimodal at mid-tier price

90
1481148 t/s$0.75 / $3.751M40.0Jul 2026
37

OpenAI · Long context

89
1310120 t/s$2 / $81M17.8Apr 2025
38

DeepSeek · Cheap frontier-adjacent tokens (GA build)

89
146162 t/s$0.66 / $1.981M67.4Aug 2026
39

MiniMax · Open-weight agentic coding

89
145580 t/s$0.6 / $2.41M59.3Jun 2026
40

OpenAI · Cheapest current GPT-6 for high-volume work

89
——$0.1 / $0.51M296.7Sep 2026
41

OpenAI · Reasoning & math

88
1305155 t/s$1.1 / $4.4200K32.0Jan 2025
42

Anthropic · Coding & balance

88
132095 t/s$3 / $15200K9.8May 2025
43

Z.ai (Zhipu AI) · Open-weight agentic & tool use

88
146748 t/s$0.98 / $3.08200K43.3Apr 2026
44

xAI · Real-time info

87
133082 t/s$3 / $15131K9.7Feb 2025
45

DeepSeek · Best open-source value

86
131062 t/s$0.27 / $1.1128K125.5Mar 2025
46

Alibaba Cloud · Multilingual & APAC

86
1448124 t/s$1.4 / $5.6256K24.6Apr 2026
47

OpenAI · General purpose

85
1285109 t/s$2.5 / $10128K13.6May 2024
48

DeepSeek · Cheapest tokens, period

84
1441105 t/s$0.14 / $0.281M400.0Apr 2026
49

OpenAI · High-volume cheap frontier-family

83
1436186 t/s$0.2 / $1.21M118.6Jul 2026
50

Meta · Open-source value

80
1260135 t/s$0.2 / $0.61M200.0Apr 2025
51

Alibaba Cloud · Open-source flagship

80
125585 t/s$0.3 / $0.9131K133.3Sep 2024
52

Alibaba Cloud · Locally-runnable open-weight (~17GB VRAM at 4-bit)

80
——Self-host262K, Aug 2026
53

Mistral AI · Multilingual

79
125078 t/s$2 / $6128K19.8Nov 2024
54

Google · High-volume cheap multimodal

79
1418192 t/s$0.3 / $2.51M56.4Jul 2026
55

NVIDIA · Efficient single-GPU deployment, 1M context

79
——Self-host1M, Aug 2026
56

xAI · Budget reasoning

78
1275165 t/s$0.3 / $0.5131K195.0Feb 2025
57

Perplexity · Search + citations

78
—65 t/s$3 / $15200K8.7Feb 2025
58

Mistral AI · Code generation

76
—195 t/s$0.3 / $0.9256K126.7Jan 2025
59

NVIDIA · Open multimodal

76
1361158 t/sSelf-host256K, Apr 2026
60

Tencent · Long-context retrieval & agentic search, no license restrictions

76
——Self-host256K, Jul 2026
61

Anthropic · Speed & cost

75
1230172 t/s$0.8 / $4200K31.3Oct 2024
62

Google · Self-hosted general purpose

75
1351142 t/sSelf-host128K, Apr 2026
63

Google · Fastest + cheapest

74
1240244 t/s$0.1 / $0.41M296.0Feb 2025
64

Alibaba Cloud · Open-source coding

74
—125 t/s$0.15 / $0.45131K246.7Nov 2024
65

Ant Group · Efficient MoE alternative to trillion-param flagships

74
——Self-host256K, Aug 2026
66

OpenAI · High throughput

72
1216183 t/s$0.15 / $0.6128K192.0Jul 2024
67

Meta · Longest context

71
1195198 t/s$0.15 / $0.410M258.2Apr 2025
68

Amazon · AWS ecosystem

70
—110 t/s$0.8 / $3.2300K35.0Dec 2024
69

Cohere · Enterprise RAG

68
117072 t/s$2.5 / $10128K10.9Aug 2024
Quality = composite benchmark (MMLU, HumanEval, MATH)Arena ELO = LMSYS Chatbot Arena ratingValue = quality per dollarPrice = input / output per 1M tokens

How We Rank AI Models

Our leaderboard uses a composite quality index that combines three key benchmarks: MMLU Pro (measuring knowledge and reasoning across 57 subjects), HumanEval (measuring code generation ability), and MATH (measuring mathematical problem-solving). Scores are normalized to a 0-100 scale and cross-referenced against LMSYS Chatbot Arena ELO ratings for real-world validation.

We track speed (tokens per second), time-to-first-token (TTFT), pricing, and context window size to give you a complete picture. The Value Score divides quality by cost, showing you which models deliver the most capability per dollar.

Key Trends in AI Model Performance

  • A new frontier #1: Anthropic's Claude Opus 5 (24 July 2026) tops the board — a step change over Opus 4.8 on deep reasoning and long-horizon agentic work, at unchanged $5/$25 pricing and roughly half the cost of Claude Fable 5. Thinking is on by default and the prompt-cache minimum halves to 512 tokens
  • OpenAI split GPT-5.6 into three tiers: Sol ($5/$30), Terra ($2/$12), and Luna ($0.20/$1.20), GA 9 July. On 30 July OpenAI cut Terra 20% and Luna 80%, citing inference work that reduced serving cost 20% — the steepest cut of the year from a US lab, and a direct answer to the Chinese open-weight tier on price
  • Open weights have caught the frontier: Moonshot's Kimi K3 (2.8T MoE) ranked #3 overall on the Artificial Analysis Intelligence Index at its July 2026 launch — ahead of every proprietary model except Claude Fable 5 and GPT-5.6 Sol at the time — and took #1 on the Frontend Code Arena. It is not a lone outlier: DeepSeek V4 Pro posts 80.6% on SWE-bench Verified (matching Gemini 3.1 Pro) and MiniMax M3 80.5%, while GLM-5.2 scores 62.1% on SWE-bench Pro against GPT-5.5's 58.6% — an MIT-licensed model you can self-host outscoring a $5/$30 US flagship on agentic coding
  • "Open" now means datacentre-scale: the leading open models are trillion-parameter MoEs — GLM-5.2 needs ~1TB VRAM in BF16 (~8x H200 at FP8), Kimi K3 needs 64+ accelerators. The ecosystem has bifurcated into giant open MoEs and genuinely small models (Qwen 3.6 27B, Gemma 4, Nemotron 3 Nano Omni) that run on one GPU or a phone, with little in between
  • Meta left the open frontier: Llama 5 has not shipped and is now forecast for 2027. Meta pivoted to its first closed frontier model, Muse Spark (April 2026), with no weights and no architecture paper — leaving Chinese labs (DeepSeek, Moonshot, Z.ai, Alibaba, MiniMax) as the effective owners of the open-weight frontier
  • Reasoning is built in: Frontier models like GPT-5.6 Sol, Claude Opus 5, and Gemini 3.1 Pro ship extended thinking by default, while deep-research variants trade latency for accuracy on the hardest tasks
  • Million-token context is standard: 1M-token windows are now table stakes for flagship models across OpenAI, Anthropic, Google, and xAI
  • DeepSeek is the price floor: V4 Pro sits at $0.435/$0.87 per 1M tokens — launched at $1.74/$3.48, cut 75% as a promotion, then made permanent on 31 May 2026 — and V4 Flash at $0.14/$0.28 is the cheapest usable model anywhere, with cache hits at $0.0028. Measured cost-per-task is about $0.04 on V4 Pro against $0.94 on Kimi K3 and far more on the closed frontier. One dollar buys ~1.15M output tokens on V4 Pro, ~227K on GLM-5.2, ~67K on Kimi K3. Watch one thing: DeepSeek has announced 2x peak-hour pricing for two daily windows, not yet active as of 4 August

Choosing the Right Model

There is no single "best" model: it depends on your use case. For most applications, a model routing approach works best: route simple queries to fast, cheap models and complex queries to frontier models. This gives you the best of both worlds: low cost and high quality.