Updated Aug 4, 2026

AI Models Leaderboard, August 2026

Large language models ranked by LMSys Arena Elo, MMLU, HumanEval, MATH, pricing, and inference speed. Refreshed monthly with live data from official provider pricing pages, Artificial Analysis, and the Arena.

In short:As of August 2026, Claude Mythos 5 tops the AI models leaderboard at 1531 Arena Elo, across 377+ models compared on quality, cost, and speed. Per-task leaders differ: check the column for your use case (text, coding, reasoning, or value).

Which AI model should you use in August 2026?

The leaderboard below ranks 377 frontier and frontier-adjacent models on a single normalized quality index, but the practical answer is task-dependent: pick the leading code/agentic model for coding, the leading long-context model for science and retrieval, the leading chat model for general use, and the open-weights leader when budget matters. Multi-model routing is now the default architectural choice for serious production deployments.

377 models
#ModelQualityArena ELOSpeedPriceContextValueReleased
1

Anthropic · Frontier agentic coding & knowledge work

100
152558 t/s$10 / $501M3.3Jun 2026
2

Anthropic · Ceiling capability (limited access)

100
153156 t/s$10 / $501M3.3Jun 2026
3

Anthropic · Frontier agentic coding & reasoning

99
152274 t/s$5 / $251M6.6Jul 2026
4

Anthropic · Coding, agents & computer use

98
151272 t/s$5 / $251M6.5May 2026
5

OpenAI · General frontier reasoning & tools

98
151496 t/s$5 / $301M5.6Jul 2026
6

OpenAI · Frontier general purpose

97
150670 t/s$5 / $301M5.5Apr 2026
7

OpenAI · Complex analysis

97
$30 / $1801M0.9Mar 2026
8

OpenAI · Complex analysis

97
$21 / $168400K1.0Dec 2025
9

Anthropic · Complex analysis

97
$30 / $1501M1.1May 2026
10

Moonshot AI · Open-weight coding frontier

97
150055 t/s$3 / $151M10.8Jul 2026
11

OpenAI · Reasoning at any cost

96
151068 t/s$30 / $1801M0.9Apr 2026
12

Anthropic · Coding & agentic workflows

96
150568 t/s$5 / $251M6.4Apr 2026
13

OpenAI · Deep research

96
$10 / $40200K3.8Oct 2025
14

OpenAI · Deep research

96
$2 / $8200K19.2Oct 2025
15

OpenAI · Hard reasoning

96
$20 / $80200K1.9Jun 2025
16

Google · Speed & cost

96
1505$2 / $121M13.7Feb 2026
17

Google · Science & long-context

96
1505131 t/s$2 / $121M13.7Apr 2026
18

Anthropic · General purpose

95
1490$5 / $251M6.3Feb 2026
19

Anthropic · General purpose

95
$5 / $25200K6.3Nov 2025
20

Anthropic · Complex analysis

95
$30 / $1501M1.1Apr 2026
21

Alibaba Cloud · Multimodal APAC frontier (unverified)

95
149686 t/s$2.5 / $7.51M19.0Jul 2026
22

Google · Image generation

94
$2 / $1266K13.4Nov 2025
23

Anthropic · Multimodal

94
$15 / $75200K2.1Aug 2025
24

OpenAI · Hard reasoning

94
137068 t/s$10 / $40200K3.8Apr 2025
25

Alibaba Cloud · Long autonomous agentic runs

94
148890 t/s$2.5 / $7.51M18.8May 2026
26

Alibaba Cloud · Long autonomous agentic runs

94
148890 t/s$2.5 / $7.51M18.8May 2026
27

xAI · Cheap output tokens & real-time info

94
149988 t/s$2 / $6500K23.5Jul 2026
28

xAI · Agentic tasks & real-time info

93
149683 t/s$1.25 / $2.51M49.6May 2026
29

OpenAI · General purpose

93
1495$2.5 / $151M10.6Mar 2026
30

OpenAI · General purpose

93
$1.75 / $14128K11.8Mar 2026
31

OpenAI · Code generation

93
$1.75 / $14400K11.8Feb 2026
32

OpenAI · Code generation

93
$1.75 / $14400K11.8Jan 2026
33

OpenAI · General purpose

93
$1.75 / $14128K11.8Dec 2025
34

OpenAI · General purpose

93
$1.75 / $14400K11.8Dec 2025
35

OpenAI · Code generation

93
$1.25 / $10400K16.5Dec 2025
36

OpenAI · General purpose

93
$1.25 / $10400K16.5Nov 2025
37

OpenAI · General purpose

93
$1.25 / $10128K16.5Nov 2025
38

OpenAI · Code generation

93
$1.25 / $10400K16.5Nov 2025
39

OpenAI · Hard reasoning

93
$150 / $600200K0.2Mar 2025
40

OpenAI · Complex analysis

93
$30 / $608K2.1May 2023
41

OpenAI · Multimodal

93
$30 / $608K2.1May 2023
42

xAI · General purpose

93
1496$1.25 / $2.52M49.6Mar 2026
43

OpenAI · Complex analysis

93
$8 / $15272K8.1Apr 2026
44

Anthropic · Balanced agents & coding value

93
147998 t/s$3 / $151M10.3Jun 2026
45

OpenAI · Default production tier

93
1489118 t/s$2 / $121M13.3Jul 2026
46

Moonshot AI · Frontier quality at low cost

92
146648 t/s$0.73 / $3.49256K43.6Apr 2026
47

Google · Multimodal + value

92
134587 t/s$1.25 / $101M16.4Mar 2025
48

Anthropic · Complex analysis

91
136052 t/s$15 / $75200K2.0May 2025
49

· Hard reasoning

91
$0.3 / $1.1164K130.0Jul 2025
50

Google · Speed & cost

91
$1.25 / $101M16.2Jun 2025
51

DeepSeek · Hard reasoning

91
$0.5 / $2.15164K68.7May 2025
52

Google · Speed & cost

91
$1.25 / $101M16.2May 2025
53

DeepSeek · Hard reasoning

91
$0.29 / $0.2933K313.8Jan 2025
54

DeepSeek · Hard reasoning

91
$0.7 / $0.8131K121.3Jan 2025
55

DeepSeek · Cheap reasoning

91
135035 t/s$0.55 / $2.19128K66.4Jan 2025
56

Moonshot AI · Open-weight agentic coding

91
55 t/s$0.73 / $3.49256K43.1Jun 2026
57

· Open-weight reasoning & tool use

91
50 t/s$0.2 / $0.8262K182.0Jun 2026
58

Z.ai (Zhipu AI) · Fastest self-hostable frontier model

91
1483168 t/s$1.4 / $4.41M31.4Jun 2026
59

DeepSeek · Cheapest frontier-class tokens

90
146762 t/s$0.435 / $0.871M137.9Apr 2026
60

Anthropic · Coding & balance

90
146773 t/s$3 / $151M10.0Feb 2026
61

OpenAI · General purpose

90
1455$1.25 / $10400K16.0Aug 2025
62

xAI · General purpose

90
$3 / $15131K10.0Apr 2025
63

Alibaba Cloud · Open-source

90
$1.04 / $6.24262K24.7Apr 2026
64

Google · Fast multimodal at mid-tier price

90
1481148 t/s$1.5 / $7.51M20.0Jul 2026
65

OpenAI · Long context

89
1310120 t/s$2 / $81M17.8Apr 2025
66

Moonshot AI · Speed & cost

89
1452$0.4 / $1.9262K77.4Jan 2026
67

MiniMax · Open-weight agentic coding

89
145580 t/s$0.6 / $2.41M59.3Jun 2026
68

· Open-weight agentic coding (provisional)

89
$0.98 / $3.08200K43.8Jun 2026
69

· Open-weight agentic & tool use

88
146748 t/s$0.98 / $3.08200K43.3Apr 2026
70

OpenAI · Multimodal

88
$10 / $10400K8.8Oct 2025
71

OpenAI · Complex analysis

88
$15 / $120400K1.3Oct 2025
72

Anthropic · General purpose

88
$3 / $151M9.8Sep 2025
73

OpenAI · General purpose

88
$2.5 / $10128K14.1Aug 2025
74

OpenAI · Search + citations

88
$2.5 / $10128K14.1Mar 2025
75

OpenAI · Hard reasoning

88
$15 / $60200K2.3Dec 2024
76

OpenAI · General purpose

88
$2.5 / $10128K14.1Nov 2024
77

OpenAI · Multimodal

88
$6 / $18128K7.3May 2024
78

OpenAI · General purpose

88
$5 / $15128K8.8May 2024
79

OpenAI · Multimodal

88
$10 / $30128K4.4Apr 2024
80

OpenAI · Complex analysis

88
$10 / $30128K4.4Jan 2024
81

OpenAI · Multimodal

88
$10 / $30128K4.4Nov 2023
82

· Open-source

88
1450$0.6 / $1.9280K69.8Feb 2026
83

Anthropic · Coding & balance

88
132095 t/s$3 / $15200K9.8May 2025
84

OpenAI · Reasoning & math

88
1305155 t/s$1.1 / $4.4200K32.0Jan 2025
85

Z.ai (Zhipu AI) · Open-weight agentic & tool use

88
146748 t/s$0.98 / $3.08200K43.3Apr 2026
86

xAI · Real-time info

87
133082 t/s$3 / $15131K9.7Feb 2025
87

DeepSeek · Open-source

87
1455$0.252 / $0.378164K276.2Dec 2025
88

· Open-source

86
$0.135 / $0.5131K270.9Dec 2025
89

DeepSeek · Open-source

86
$0.287 / $0.431164K239.6Dec 2025
90

DeepSeek · Open-source

86
$0.27 / $0.41164K252.9Sep 2025
91

DeepSeek · Open-source

86
$0.27 / $0.95164K141.0Sep 2025
92

DeepSeek · Open-source

86
$0.21 / $0.7933K172.0Aug 2025
93

DeepSeek · Open-source

86
$0.2 / $0.77164K177.3Mar 2025
94

Anthropic · General purpose

86
$3 / $15200K9.6Feb 2025
95

Anthropic · Hard reasoning

86
$3 / $15200K9.6Feb 2025
96

DeepSeek · Best open-source value

86
131062 t/s$0.27 / $1.1128K125.5Mar 2025
97

Alibaba Cloud · Multilingual & APAC

86
1448124 t/s$1.4 / $5.6256K24.6Apr 2026
98

Alibaba Cloud · Multilingual & APAC

86
1448124 t/s$1.4 / $5.6256K24.6Apr 2026
99

OpenAI · General purpose

85
1285109 t/s$2.5 / $10128K13.6May 2024
100

OpenAI · General purpose

85
1285109 t/s$2.5 / $10128K13.6May 2024
Page 1 of 4 · 1100 of 377
Quality = composite benchmark (MMLU, HumanEval, MATH)Arena ELO = LMSYS Chatbot Arena ratingValue = quality per dollarPrice = input / output per 1M tokens

How the LLM leaderboard works

We pull official provider pricing every 24 hours, Artificial Analysis benchmark snapshots weekly, and LMSys Arena Elo as it publishes. The composite quality index is a 0-100 normalization over MMLU Pro, HumanEval, and MATH, weighted by recency and cross-validated against Arena Elo. We do not accept vendor-supplied numbers without an independent reference.

Where the leaderboard is wrong

No leaderboard predicts your production accuracy. LMSys Arena rewards style and short-conversation polish; a top-Arena model can still under-perform on your specific function-calling schema or long-context retrieval workload. Build an internal eval harness before you commit. See our LMArena Elo explained and LLM routing writeups for the deep-dive.

Related rankings