Updated Sep 21, 2026

LMArena.ai, Top Models September 2026

LMArena.ai is the rebranded LMSys Chatbot Arena. Same blind pairwise-voting methodology, same Elo math, new home. Here is who leads each board this month and what the rebrand actually changed for buyers.

The LMSys to LMArena.ai story

The Chatbot Arena began in 2023 as a research project under LMSys, an academic group out of UC Berkeley. It quickly became the most-cited LLM benchmark because it measured something the capability-only benchmarks could not: actual human preference under blind side-by-side comparison. By 2024 the project had processed millions of votes, become a procurement input for Fortune 500 buyers, and outgrown its original academic scaffold. The 2024-25 transition to the lmarena.ai domain consolidated the project as an independent organisation while keeping the same Elo methodology and open vote pool.

For users the rebrand changed almost nothing: same prompts, same blind voting, same Elo math. For procurement teams the rebrand codified Arena Elo as a vendor-neutral signal independent of any single university. That is what made it sticky as a reference point in enterprise contracts.

The four-way race at the top

The September 2026 snapshot below shows Claude Fable 5 and the newly released Claude Opus 5 (24 July) at the top of the text board, with GPT-5.6 Sol and Grok 4.5 pressing close behind. The structural story is underneath them: Moonshot's Kimi K3 took #1 on the coding (Frontend Code) Arena at 1,679 Elo — the first open model to top a board outright — and sits at ~1500 on text, level with the proprietary pack. Below it, GLM-5.2 and DeepSeek V4 Pro put two more MIT-licensed models inside the frontier Elo band. The top of LMArena.ai is no longer a closed-vendor race.

Top of LMArena.ai text leaderboard (September 2026)

  Claude Fable 5           1525   ████████████████████   text #1
  Claude Opus 5            1522   ████████████████████   new flagship (24 Jul)
  GPT-5.6 Sol              1514   ███████████████████    OpenAI flagship tier
  Claude Opus 4.8          1512   ███████████████████    prior text #1
  Grok 4.5                 1499   ███████████████████    cheapest frontier output
  Gemini 3.1 Pro Preview   1500   ███████████████████    science leader
  Kimi K3 (open)           1500   ███████████████████    coding #1 · open weight
  GPT-5.5 Pro              1488   ██████████████████     reasoning
  GLM-5.2 (open, MIT)      1483   ██████████████████     beats GPT-5.5 on SWE Pro
  Claude Sonnet 5          1479   ██████████████████     workhorse tier
  DeepSeek V4 Pro (MIT)    1462   █████████████████      $0.435/$0.87 · cheapest
  Qwen 3.7 Max             1455   ████████████████       top Chinese proprietary
  DeepSeek V4 Flash (MIT)  1441   ████████████████       $0.14/$0.28 · price floor
  Llama 4 Maverick         1352   ███████████            Meta's last open model

Full Leaderboard

56 models
#ModelQualityArena ELOSpeedPriceContextValueReleased
1

Anthropic · Frontier agentic coding & knowledge work

100
152558 t/s$10 / $501M3.3Jun 2026
2

Anthropic · Ceiling capability (limited access)

100
153156 t/s$10 / $501M3.3Jun 2026
3

Anthropic · Frontier agentic coding & reasoning

99
152274 t/s$5 / $251M6.6Jul 2026
4

Anthropic · Coding, agents & computer use

98
151272 t/s$5 / $251M6.5May 2026
5

OpenAI · General frontier reasoning & tools

98
151496 t/s$5 / $301M5.6Jul 2026
6

OpenAI · Frontier general purpose

97
150670 t/s$5 / $301M5.5Apr 2026
7

Moonshot AI · Open-weight coding frontier

97
150055 t/s$3 / $151M10.8Jul 2026
8

OpenAI · Reasoning at any cost

96
151068 t/s$30 / $1801M0.9Apr 2026
9

Anthropic · Coding & agentic workflows

96
150568 t/s$5 / $251M6.4Apr 2026
10

Google · Science & long-context

96
1505131 t/s$2 / $121M13.7Apr 2026
11

Alibaba Cloud · Multimodal APAC frontier (hosted); text-only open weights

96
149186 t/s$2 / $61M24.0Aug 2026
12

OpenAI · Hard reasoning

94
137068 t/s$2 / $8200K18.8Apr 2025
13

Alibaba Cloud · Long autonomous agentic runs

94
148890 t/s$2.5 / $7.51M18.8May 2026
14

xAI · Cheap output tokens & real-time info

94
149988 t/s$2 / $6500K23.5Jul 2026
15

xAI · Agentic tasks & real-time info

93
149683 t/s$1.25 / $2.51M49.6May 2026
16

OpenAI · Default production tier

93
1489118 t/s$2 / $121M13.3Jul 2026
17

Anthropic · Balanced agents & coding value

93
147998 t/s$3 / $151M10.3Jun 2026
18

Google · Multimodal + value

92
134587 t/s$1.25 / $101M16.4Mar 2025
19

Moonshot AI · Frontier quality at low cost

92
146648 t/s$0.73 / $3.49256K43.6Apr 2026
20

Anthropic · Complex analysis

91
136052 t/s$15 / $75200K2.0May 2025
21

DeepSeek · Cheap reasoning

91
135035 t/s$0.55 / $2.19128K66.4Jan 2025
22

Z.ai (Zhipu AI) · Fastest self-hostable frontier model

91
1483168 t/s$1.4 / $4.41M31.4Jun 2026
23

Anthropic · Coding & balance

90
146773 t/s$3 / $151M10.0Feb 2026
24

Google · Fast multimodal at mid-tier price

90
1481148 t/s$0.75 / $3.751M40.0Jul 2026
25

OpenAI · Long context

89
1310120 t/s$2 / $81M17.8Apr 2025
26

DeepSeek · Cheap frontier-adjacent tokens (GA build)

89
146162 t/s$0.66 / $1.981M67.4Aug 2026
27

MiniMax · Open-weight agentic coding

89
145580 t/s$0.6 / $2.41M59.3Jun 2026
28

OpenAI · Reasoning & math

88
1305155 t/s$1.1 / $4.4200K32.0Jan 2025
29

Anthropic · Coding & balance

88
132095 t/s$3 / $15200K9.8May 2025
30

Z.ai (Zhipu AI) · Open-weight agentic & tool use

88
146748 t/s$0.98 / $3.08200K43.3Apr 2026
31

xAI · Real-time info

87
133082 t/s$3 / $15131K9.7Feb 2025
32

DeepSeek · Best open-source value

86
131062 t/s$0.27 / $1.1128K125.5Mar 2025
33

Alibaba Cloud · Multilingual & APAC

86
1448124 t/s$1.4 / $5.6256K24.6Apr 2026
34

OpenAI · General purpose

85
1285109 t/s$2.5 / $10128K13.6May 2024
35

DeepSeek · Cheapest tokens, period

84
1441105 t/s$0.14 / $0.281M400.0Apr 2026
36

OpenAI · High-volume cheap frontier-family

83
1436186 t/s$0.2 / $1.21M118.6Jul 2026
37

Meta · Open-source value

80
1260135 t/s$0.2 / $0.61M200.0Apr 2025
38

Alibaba Cloud · Open-source flagship

80
125585 t/s$0.3 / $0.9131K133.3Sep 2024
39

Alibaba Cloud · Locally-runnable open-weight (~17GB VRAM at 4-bit)

80
Self-host262K, Aug 2026
40

Mistral AI · Multilingual

79
125078 t/s$2 / $6128K19.8Nov 2024
41

Google · High-volume cheap multimodal

79
1418192 t/s$0.3 / $2.51M56.4Jul 2026
42

NVIDIA · Efficient single-GPU deployment, 1M context

79
Self-host1M, Aug 2026
43

xAI · Budget reasoning

78
1275165 t/s$0.3 / $0.5131K195.0Feb 2025
44

Perplexity · Search + citations

78
65 t/s$3 / $15200K8.7Feb 2025
45

Mistral AI · Code generation

76
195 t/s$0.3 / $0.9256K126.7Jan 2025
46

NVIDIA · Open multimodal

76
1361158 t/sSelf-host256K, Apr 2026
47

Tencent · Long-context retrieval & agentic search, no license restrictions

76
Self-host256K, Jul 2026
48

Anthropic · Speed & cost

75
1230172 t/s$0.8 / $4200K31.3Oct 2024
49

Google · Self-hosted general purpose

75
1351142 t/sSelf-host128K, Apr 2026
50

Google · Fastest + cheapest

74
1240244 t/s$0.1 / $0.41M296.0Feb 2025
51

Alibaba Cloud · Open-source coding

74
125 t/s$0.15 / $0.45131K246.7Nov 2024
52

Ant Group · Efficient MoE alternative to trillion-param flagships

74
Self-host256K, Aug 2026
53

OpenAI · High throughput

72
1216183 t/s$0.15 / $0.6128K192.0Jul 2024
54

Meta · Longest context

71
1195198 t/s$0.15 / $0.410M258.2Apr 2025
55

Amazon · AWS ecosystem

70
110 t/s$0.8 / $3.2300K35.0Dec 2024
56

Cohere · Enterprise RAG

68
117072 t/s$2.5 / $10128K10.9Aug 2024
Quality = composite benchmark (MMLU, HumanEval, MATH)Arena ELO = LMSYS Chatbot Arena ratingValue = quality per dollarPrice = input / output per 1M tokens

What to do this quarter

  1. Update bookmarks and citations. Internal eval-spec docs and procurement RFPs that reference "lmsys.org" should be updated to lmarena.ai. The data continues at the new domain.
  2. Pull from the right board. Coding teams should cite the coding Arena Elo (Moonshot's Kimi K3 leads the Frontend Code Arena at ~1,679, ahead of Claude Fable 5, GPT-5.6 Sol, and Opus 4.8). Generic chat teams should cite the text leaderboard (Claude Fable 5 and Opus 5 lead at ~1522-1525, with GPT-5.6 Sol, Gemini 3.1 Pro Preview, and Kimi K3 clustered near 1500).
  3. Build dual-vendor capability. The top four models are within 40 Elo of each other. Treat them as interchangeable on capability and optimise for switching cost.
  4. Pair Arena scores with workload-specific evals. Arena rewards short-conversation polish. Long-context, tool-use, and domain-specific tasks need their own measurement.
  5. Track the open-weight gap. DeepSeek V4 Pro under Apache 2.0 sits at 1462 Elo, within 38 points of the text leader. The gap is the smallest it has ever been.
  6. Watch GPT-5.5 Pro pricing. At $30/$180 per 1M tokens, paying for the top of LMArena.ai now costs 200x more per token than the cheapest tier. The cost curve is steepening.
  7. Re-baseline at every model launch. Tokenizer changes between releases (as seen across the Claude Opus 4.6 → 4.7 → 4.8 line) shift effective cost without shifting list price.

Related reading

Teams running side-by-side evals against multiple LMArena.ai leaders typically expose them through Swfte Connect as a single endpoint, then run their own internal Elo on production prompts. That is the only way to verify whether public Arena rank translates to your workload.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.