Updated Jul 9, 2026

LMArena.ai — Top Models July 2026

LMArena.ai is the rebranded LMSys Chatbot Arena. Same blind pairwise-voting methodology, same Elo math, new home. Here is who leads each board this month and what the rebrand actually changed for buyers.

The LMSys to LMArena.ai story

The Chatbot Arena began in 2023 as a research project under LMSys, an academic group out of UC Berkeley. It quickly became the most-cited LLM benchmark because it measured something the capability-only benchmarks could not: actual human preference under blind side-by-side comparison. By 2024 the project had processed millions of votes, become a procurement input for Fortune 500 buyers, and outgrown its original academic scaffold. The 2024-25 transition to the lmarena.ai domain consolidated the project as an independent organisation while keeping the same Elo methodology and open vote pool.

For users the rebrand changed almost nothing: same prompts, same blind voting, same Elo math. For procurement teams the rebrand codified Arena Elo as a vendor-neutral signal independent of any single university. That is what made it sticky as a reference point in enterprise contracts.

The four-way race at the top

The July 2026 snapshot below shows Claude Opus 4.8 as the new text and coding leader, with Gemini 3.1 Pro and Opus 4.7 just behind at the historical 1500 Elo barrier. The top of LMArena.ai is now genuinely contested.

Top of LMArena.ai text leaderboard (July 2026)

  Claude Opus 4.8 Thinking 1512   ████████████████████   text & coding #1
  Gemini 3.1 Pro Preview   1500   ███████████████████    science leader
  Claude Opus 4.7 Thinking 1495   ███████████████████    prior coding #1
  GPT-5.5 Pro              1488   ██████████████████     reasoning
  DeepSeek V4 Pro          1462   █████████████████      Apache 2.0
  Qwen 3.7 Max             1455   ████████████████       top Chinese model
  Claude Sonnet 4          1402   ██████████████         workhorse tier
  GPT-4.1                  1395   █████████████          legacy frontier
  Gemini 2.5 Pro           1388   █████████████          legacy frontier
  Llama 4 Maverick         1352   ███████████            open weights
  Mistral Large 3          1341   ███████████            open weights

Full Leaderboard

362 models
#ModelQualityArena ELOSpeedPriceContextValueReleased
1

Anthropic · Frontier agentic coding & knowledge work

100
152558 t/s$10 / $501M3.3Jun 2026
2

Anthropic · Coding, agents & computer use

99
151272 t/s$5 / $251M6.6May 2026
3

OpenAI · Reasoning at any cost

98
151068 t/s$30 / $1801M0.9Apr 2026
4

OpenAI · General frontier reasoning & tools

98
151496 t/s$5 / $30400K5.6Jul 2026
5

OpenAI · Frontier general purpose

97
150670 t/s$5 / $301M5.5Apr 2026
6

OpenAI · Complex analysis

97
$30 / $1801M0.9Mar 2026
7

OpenAI · Complex analysis

97
$21 / $168400K1.0Dec 2025
8

Anthropic · Complex analysis

97
$30 / $1501M1.1May 2026
9

Google · Long-context & multimodal

97
1508128 t/s$2 / $122M13.9Jul 2026
10

Anthropic · Coding & agentic workflows

96
150568 t/s$5 / $251M6.4Apr 2026
11

OpenAI · Deep research

96
$10 / $40200K3.8Oct 2025
12

OpenAI · Deep research

96
$2 / $8200K19.2Oct 2025
13

OpenAI · Hard reasoning

96
$20 / $80200K1.9Jun 2025
14

Google · Speed & cost

96
1505$2 / $121M13.7Feb 2026
15

Google · Science & long-context

96
1505131 t/s$2 / $121M13.7Apr 2026
16

Anthropic · General purpose

95
1490$5 / $251M6.3Feb 2026
17

Anthropic · General purpose

95
$5 / $25200K6.3Nov 2025
18

Anthropic · Complex analysis

95
$30 / $1501M1.1Apr 2026
19

Google · Image generation

94
$2 / $1266K13.4Nov 2025
20

Anthropic · Multimodal

94
$15 / $75200K2.1Aug 2025
21

OpenAI · Hard reasoning

94
137068 t/s$10 / $40200K3.8Apr 2025
22

Alibaba Cloud · Long autonomous agentic runs

94
148890 t/s$2.5 / $7.51M18.8May 2026
23

xAI · Agentic tasks & real-time info

93
149683 t/s$1.25 / $2.51M49.6May 2026
24

OpenAI · General purpose

93
1495$2.5 / $151M10.6Mar 2026
25

OpenAI · General purpose

93
$1.75 / $14128K11.8Mar 2026
26

OpenAI · Code generation

93
$1.75 / $14400K11.8Feb 2026
27

OpenAI · Code generation

93
$1.75 / $14400K11.8Jan 2026
28

OpenAI · General purpose

93
$1.75 / $14128K11.8Dec 2025
29

OpenAI · General purpose

93
$1.75 / $14400K11.8Dec 2025
30

OpenAI · Code generation

93
$1.25 / $10400K16.5Dec 2025
31

OpenAI · General purpose

93
$1.25 / $10400K16.5Nov 2025
32

OpenAI · General purpose

93
$1.25 / $10128K16.5Nov 2025
33

OpenAI · Code generation

93
$1.25 / $10400K16.5Nov 2025
34

OpenAI · Hard reasoning

93
$150 / $600200K0.2Mar 2025
35

OpenAI · Complex analysis

93
$30 / $608K2.1May 2023
36

OpenAI · Multimodal

93
$30 / $608K2.1May 2023
37

xAI · General purpose

93
1496$1.25 / $2.52M49.6Mar 2026
38

OpenAI · Complex analysis

93
$8 / $15272K8.1Apr 2026
39

Anthropic · Balanced agents & coding value

93
147998 t/s$3 / $151M10.3Jun 2026
40

Moonshot AI · Frontier quality at low cost

92
146648 t/s$0.73 / $3.49256K43.6Apr 2026
41

Google · Multimodal + value

92
134587 t/s$1.25 / $101M16.4Mar 2025
42

DeepSeek · Open-weight reasoning value

92
147162 t/s$0.5 / $1.1256K115.0Jul 2026
43

Anthropic · Complex analysis

91
136052 t/s$15 / $75200K2.0May 2025
44

· Hard reasoning

91
$0.3 / $1.1164K130.0Jul 2025
45

Google · Speed & cost

91
$1.25 / $101M16.2Jun 2025
46

DeepSeek · Hard reasoning

91
$0.5 / $2.15164K68.7May 2025
47

Google · Speed & cost

91
$1.25 / $101M16.2May 2025
48

DeepSeek · Hard reasoning

91
$0.29 / $0.2933K313.8Jan 2025
49

DeepSeek · Hard reasoning

91
$0.7 / $0.8131K121.3Jan 2025
50

DeepSeek · Hard reasoning

91
$0.7 / $2.564K56.9Jan 2025
51

Moonshot AI · Open-weight agentic coding

91
55 t/s$0.73 / $3.49256K43.1Jun 2026
52

· Open-weight reasoning & tool use

91
50 t/s$0.2 / $0.8262K182.0Jun 2026
53

Meta · Open-weight general & multilingual

91
146676 t/s$0.8 / $2.41M56.9Jun 2026
54

DeepSeek · Open-source value leader

90
146733 t/s$1.74 / $3.481M34.5Apr 2026
55

Anthropic · Coding & balance

90
146773 t/s$3 / $151M10.0Feb 2026
56

OpenAI · General purpose

90
1455$1.25 / $10400K16.0Aug 2025
57

xAI · General purpose

90
$3 / $15131K10.0Apr 2025
58

Alibaba Cloud · Open-source

90
$1.04 / $6.24262K24.7Apr 2026
59

OpenAI · Long context

89
1310120 t/s$2 / $81M17.8Apr 2025
60

Moonshot AI · Speed & cost

89
1452$0.4 / $1.9262K77.4Jan 2026
61

MiniMax · Open-weight agentic coding

89
145580 t/s$0.6 / $2.41M59.3Jun 2026
62

· Open-weight agentic coding (provisional)

89
$0.98 / $3.08200K43.8Jun 2026
63

· Open-weight agentic & tool use

88
146748 t/s$0.98 / $3.08200K43.3Apr 2026
64

OpenAI · Multimodal

88
$10 / $10400K8.8Oct 2025
65

OpenAI · Complex analysis

88
$15 / $120400K1.3Oct 2025
66

Anthropic · General purpose

88
$3 / $151M9.8Sep 2025
67

OpenAI · General purpose

88
$2.5 / $10128K14.1Aug 2025
68

OpenAI · Search + citations

88
$2.5 / $10128K14.1Mar 2025
69

OpenAI · Hard reasoning

88
$15 / $60200K2.3Dec 2024
70

OpenAI · General purpose

88
$2.5 / $10128K14.1Nov 2024
71

OpenAI · General purpose

88
$2.5 / $10128K14.1May 2024
72

OpenAI · Multimodal

88
$6 / $18128K7.3May 2024
73

OpenAI · General purpose

88
$5 / $15128K8.8May 2024
74

OpenAI · Multimodal

88
$10 / $30128K4.4Apr 2024
75

OpenAI · Complex analysis

88
$10 / $30128K4.4Jan 2024
76

OpenAI · Multimodal

88
$10 / $30128K4.4Nov 2023
77

· Open-source

88
1450$0.6 / $1.9280K69.8Feb 2026
78

Anthropic · Coding & balance

88
132095 t/s$3 / $15200K9.8May 2025
79

OpenAI · Reasoning & math

88
1305155 t/s$1.1 / $4.4200K32.0Jan 2025
80

xAI · Real-time info

87
133082 t/s$3 / $15131K9.7Feb 2025
81

DeepSeek · Open-source

87
1455$0.252 / $0.378164K276.2Dec 2025
82

· Open-source

86
$0.135 / $0.5131K270.9Dec 2025
83

DeepSeek · Open-source

86
$0.287 / $0.431164K239.6Dec 2025
84

DeepSeek · Open-source

86
$0.27 / $0.41164K252.9Sep 2025
85

DeepSeek · Open-source

86
$0.27 / $0.95164K141.0Sep 2025
86

DeepSeek · Open-source

86
$0.21 / $0.7933K172.0Aug 2025
87

DeepSeek · Open-source

86
$0.2 / $0.77164K177.3Mar 2025
88

Anthropic · General purpose

86
$3 / $15200K9.6Feb 2025
89

Anthropic · Hard reasoning

86
$3 / $15200K9.6Feb 2025
90

DeepSeek · Best open-source value

86
131062 t/s$0.27 / $1.1128K125.5Mar 2025
91

Alibaba Cloud · Multilingual & APAC

86
1448124 t/s$1.4 / $5.6256K24.6Apr 2026
92

OpenAI · General purpose

85
1285109 t/s$2.5 / $10128K13.6May 2024
93

Mistral AI · Open-source

85
$0.5 / $1.5262K85.0Dec 2025
94

Mistral AI · Open-source

85
$2 / $6131K21.3Nov 2024
95

Mistral AI · Open-source

85
$2 / $6128K21.3Feb 2024
96

Google · Speed & cost

84
$1.5 / $91M16.0May 2026
97

· Accessible open-weight agentics

84
110 t/s$0.05 / $0.2262K672.0Jun 2026
98

OpenAI · Speed & cost

83
$0.75 / $4.5400K31.6Mar 2026
99

OpenAI · Speed & cost

83
$0.25 / $2400K73.8Aug 2025
100

Alibaba Cloud · Open-source

82
$0.04 / $0.15256K863.2Mar 2026
Page 1 of 4 · 1100 of 362
Quality = composite benchmark (MMLU, HumanEval, MATH)Arena ELO = LMSYS Chatbot Arena ratingValue = quality per dollarPrice = input / output per 1M tokens

What to do this quarter

  1. Update bookmarks and citations. Internal eval-spec docs and procurement RFPs that reference "lmsys.org" should be updated to lmarena.ai. The data continues at the new domain.
  2. Pull from the right board. Coding teams should cite the coding Arena Elo (Claude Opus 4.8 now leads at ~1582, ahead of Opus 4.7 at 1567). Generic chat teams should cite the text leaderboard (Opus 4.8 leads at ~1512, Gemini 3.1 Pro Preview at ~1500).
  3. Build dual-vendor capability. The top four models are within 40 Elo of each other. Treat them as interchangeable on capability and optimise for switching cost.
  4. Pair Arena scores with workload-specific evals. Arena rewards short-conversation polish. Long-context, tool-use, and domain-specific tasks need their own measurement.
  5. Track the open-weight gap. DeepSeek V4 Pro under Apache 2.0 sits at 1462 Elo, within 38 points of the text leader. The gap is the smallest it has ever been.
  6. Watch GPT-5.5 Pro pricing. At $30/$180 per 1M tokens, paying for the top of LMArena.ai now costs 200x more per token than the cheapest tier. The cost curve is steepening.
  7. Re-baseline at every model launch. Tokenizer changes between releases (as seen across the Claude Opus 4.6 → 4.7 → 4.8 line) shift effective cost without shifting list price.

Related reading

Teams running side-by-side evals against multiple LMArena.ai leaders typically expose them through Swfte Connect as a single endpoint, then run their own internal Elo on production prompts. That is the only way to verify whether public Arena rank translates to your workload.