Updated Jul 10, 2026

AI Model Leaderboard — July 2026

Every major AI model ranked by quality, speed, pricing, and value. Filter by category, sort by any metric, and find the right model for your use case. Live data refreshed regularly with LMSys Arena Elo, official provider pricing, and Artificial Analysis benchmarks.

In short:As of July 2026, Claude Fable 5 leads the AI model leaderboard at 100/100 on our composite quality index, across 45+ models ranked by quality, price, speed, and value. The best pick depends on the workload — sort the table by the metric that matters to you.

🥇

Claude Fable 5

100/100

🥈

Claude Opus 4.8

99/100

🥉

GPT-5.5 Pro

98/100

Climb the Leaderboard

Stop reading — start ranking

Monthly Snapshot

July 2026: Top Models, Best Value, Fastest Inference

The July 2026 ranking covers 45 models across LMSys Arena Elo, MMLU Pro, HumanEval, MATH, pricing, and inference speed. Top of the table: Claude Fable 5 at 100/100 quality. The full table below is sortable by any metric. Live data is refreshed regularly from official provider pricing pages and the public Arena.

Top 5 by Quality Index

  1. Claude Fable 5 100/100
  2. Claude Opus 4.8 99/100
  3. GPT-5.5 Pro 98/100
  4. GPT-5.6 98/100
  5. GPT-5.5 97/100

Best Price-to-Quality

  1. DeepSeek V4 Flash — $0.2/1M out
  2. Gemini 2.0 Flash — $0.4/1M out
  3. Llama 4 Scout — $0.4/1M out
  4. Qwen 2.5 Coder 32B — $0.45/1M out
  5. Grok 3 Mini — $0.5/1M out

See our LMSys Arena deep dive and the monthly release roundup.

45 models
#ModelQualityArena ELOSpeedPriceContextValueReleased
1

Anthropic · Frontier agentic coding & knowledge work

100
152558 t/s$10 / $501M3.3Jun 2026
2

Anthropic · Coding, agents & computer use

99
151272 t/s$5 / $251M6.6May 2026
3

OpenAI · Reasoning at any cost

98
151068 t/s$30 / $1801M0.9Apr 2026
4

OpenAI · General frontier reasoning & tools

98
151496 t/s$5 / $30400K5.6Jul 2026
5

OpenAI · Frontier general purpose

97
150670 t/s$5 / $301M5.5Apr 2026
6

Google · Long-context & multimodal

97
1508128 t/s$2 / $122M13.9Jul 2026
7

Anthropic · Coding & agentic workflows

96
150568 t/s$5 / $251M6.4Apr 2026
8

Google · Science & long-context

96
1505131 t/s$2 / $121M13.7Apr 2026
9

OpenAI · Hard reasoning

94
137068 t/s$10 / $40200K3.8Apr 2025
10

Alibaba Cloud · Long autonomous agentic runs

94
148890 t/s$2.5 / $7.51M18.8May 2026
11

xAI · Agentic tasks & real-time info

93
149683 t/s$1.25 / $2.51M49.6May 2026
12

Anthropic · Balanced agents & coding value

93
147998 t/s$3 / $151M10.3Jun 2026
13

Google · Multimodal + value

92
134587 t/s$1.25 / $101M16.4Mar 2025
14

Moonshot AI · Frontier quality at low cost

92
146648 t/s$0.73 / $3.49256K43.6Apr 2026
15

DeepSeek · Open-weight reasoning value

92
147162 t/s$0.5 / $1.1256K115.0Jul 2026
16

Anthropic · Complex analysis

91
136052 t/s$15 / $75200K2.0May 2025
17

DeepSeek · Cheap reasoning

91
135035 t/s$0.55 / $2.19128K66.4Jan 2025
18
Llama 5 NewOSS

Meta · Open-weight general & multilingual

91
146676 t/s$0.8 / $2.41M56.9Jun 2026
19

DeepSeek · Open-source value leader

90
146733 t/s$1.74 / $3.481M34.5Apr 2026
20

Anthropic · Coding & balance

90
146773 t/s$3 / $151M10.0Feb 2026
21

OpenAI · Long context

89
1310120 t/s$2 / $81M17.8Apr 2025
22

MiniMax · Open-weight agentic coding

89
145580 t/s$0.6 / $2.41M59.3Jun 2026
23

OpenAI · Reasoning & math

88
1305155 t/s$1.1 / $4.4200K32.0Jan 2025
24

Anthropic · Coding & balance

88
132095 t/s$3 / $15200K9.8May 2025
25

Z.ai (Zhipu AI) · Open-weight agentic & tool use

88
146748 t/s$0.98 / $3.08200K43.3Apr 2026
26

xAI · Real-time info

87
133082 t/s$3 / $15131K9.7Feb 2025
27

DeepSeek · Best open-source value

86
131062 t/s$0.27 / $1.1128K125.5Mar 2025
28

Alibaba Cloud · Multilingual & APAC

86
1448124 t/s$1.4 / $5.6256K24.6Apr 2026
29

OpenAI · General purpose

85
1285109 t/s$2.5 / $10128K13.6May 2024
30

Meta · Open-source value

80
1260135 t/s$0.2 / $0.61M200.0Apr 2025
31

Alibaba Cloud · Open-source flagship

80
125585 t/s$0.3 / $0.9131K133.3Sep 2024
32

DeepSeek · Cheap-and-fast cascade tier

80
1410105 t/s$0.1 / $0.21M533.3Apr 2026
33

Mistral AI · Multilingual

79
125078 t/s$2 / $6128K19.8Nov 2024
34

xAI · Budget reasoning

78
1275165 t/s$0.3 / $0.5131K195.0Feb 2025
35

Perplexity · Search + citations

78
65 t/s$3 / $15200K8.7Feb 2025
36

Mistral AI · Code generation

76
195 t/s$0.3 / $0.9256K126.7Jan 2025
37

Mistral AI · Open multimodal

76
1361158 t/sSelf-host256KApr 2026
38

Anthropic · Speed & cost

75
1230172 t/s$0.8 / $4200K31.3Oct 2024
39

Google · Self-hosted general purpose

75
1351142 t/sSelf-host128KApr 2026
40

Google · Fastest + cheapest

74
1240244 t/s$0.1 / $0.41M296.0Feb 2025
41

Alibaba Cloud · Open-source coding

74
125 t/s$0.15 / $0.45131K246.7Nov 2024
42

OpenAI · High throughput

72
1216183 t/s$0.15 / $0.6128K192.0Jul 2024
43

Meta · Longest context

71
1195198 t/s$0.15 / $0.410M258.2Apr 2025
44

Amazon · AWS ecosystem

70
110 t/s$0.8 / $3.2300K35.0Dec 2024
45

Cohere · Enterprise RAG

68
117072 t/s$2.5 / $10128K10.9Aug 2024
Quality = composite benchmark (MMLU, HumanEval, MATH)Arena ELO = LMSYS Chatbot Arena ratingValue = quality per dollarPrice = input / output per 1M tokens

How We Rank AI Models

Our leaderboard uses a composite quality index that combines three key benchmarks: MMLU Pro (measuring knowledge and reasoning across 57 subjects), HumanEval (measuring code generation ability), and MATH (measuring mathematical problem-solving). Scores are normalized to a 0-100 scale and cross-referenced against LMSYS Chatbot Arena ELO ratings for real-world validation.

We track speed (tokens per second), time-to-first-token (TTFT), pricing, and context window size to give you a complete picture. The Value Score divides quality by cost, showing you which models deliver the most capability per dollar.

Key Trends in AI Model Performance

  • July 2026 refresh: five new models joined the board — OpenAI's GPT-5.6 (AA Index 61.0, effectively free over GPT-5.5), Google's Gemini 3.2 Pro (2M context, still $2/$12), Anthropic's Claude Sonnet 5 (Opus-class coding at a third of the price), and two open-weight releases, DeepSeek V4.5 (MIT) and Meta's Llama 5
  • A new frontier #1: Claude Opus 4.8 (May 28, 2026) tops the Artificial Analysis Intelligence Index at 61.4, edging out GPT-5.5, with standout gains in coding (SWE-bench Pro 69.2%) and computer-use agents (Online-Mind2Web 84%). Alibaba's Qwen 3.7 Max debuts as the highest-ranked Chinese model at #5
  • Open weights closing the gap: DeepSeek V4 Pro, GLM-5.1, and Kimi K2.6 now trade blows with closed frontier models on reasoning and coding — at a fraction of the price
  • Reasoning is built in: Frontier models like GPT-5.5, Claude Opus 4.8, and Gemini 3.1 Pro ship extended thinking by default, while deep-research variants trade latency for accuracy on the hardest tasks
  • Million-token context is standard: 1M-token windows are now table stakes for flagship models across OpenAI, Anthropic, Google, and xAI
  • Price keeps falling: Fast- and flash-tier models pair strong quality with low latency, while launch promos (such as DeepSeek V4) push frontier-class quality below $1 per million tokens

Choosing the Right Model

There is no single "best" model — it depends on your use case. For most applications, a model routing approach works best: route simple queries to fast, cheap models and complex queries to frontier models. This gives you the best of both worlds — low cost and high quality.