LMArena.ai, Top Models September 2026
LMArena.ai is the rebranded LMSys Chatbot Arena. Same blind pairwise-voting methodology, same Elo math, new home. Here is who leads each board this month and what the rebrand actually changed for buyers.
The LMSys to LMArena.ai story
The Chatbot Arena began in 2023 as a research project under LMSys, an academic group out of UC Berkeley. It quickly became the most-cited LLM benchmark because it measured something the capability-only benchmarks could not: actual human preference under blind side-by-side comparison. By 2024 the project had processed millions of votes, become a procurement input for Fortune 500 buyers, and outgrown its original academic scaffold. The 2024-25 transition to the lmarena.ai domain consolidated the project as an independent organisation while keeping the same Elo methodology and open vote pool.
For users the rebrand changed almost nothing: same prompts, same blind voting, same Elo math. For procurement teams the rebrand codified Arena Elo as a vendor-neutral signal independent of any single university. That is what made it sticky as a reference point in enterprise contracts.
The four-way race at the top
The September 2026 snapshot below shows Claude Fable 5 and the newly released Claude Opus 5 (24 July) at the top of the text board, with GPT-5.6 Sol and Grok 4.5 pressing close behind. The structural story is underneath them: Moonshot's Kimi K3 took #1 on the coding (Frontend Code) Arena at 1,679 Elo — the first open model to top a board outright — and sits at ~1500 on text, level with the proprietary pack. Below it, GLM-5.2 and DeepSeek V4 Pro put two more MIT-licensed models inside the frontier Elo band. The top of LMArena.ai is no longer a closed-vendor race.
Top of LMArena.ai text leaderboard (September 2026) Claude Fable 5 1525 ████████████████████ text #1 Claude Opus 5 1522 ████████████████████ new flagship (24 Jul) GPT-5.6 Sol 1514 ███████████████████ OpenAI flagship tier Claude Opus 4.8 1512 ███████████████████ prior text #1 Grok 4.5 1499 ███████████████████ cheapest frontier output Gemini 3.1 Pro Preview 1500 ███████████████████ science leader Kimi K3 (open) 1500 ███████████████████ coding #1 · open weight GPT-5.5 Pro 1488 ██████████████████ reasoning GLM-5.2 (open, MIT) 1483 ██████████████████ beats GPT-5.5 on SWE Pro Claude Sonnet 5 1479 ██████████████████ workhorse tier DeepSeek V4 Pro (MIT) 1462 █████████████████ $0.435/$0.87 · cheapest Qwen 3.7 Max 1455 ████████████████ top Chinese proprietary DeepSeek V4 Flash (MIT) 1441 ████████████████ $0.14/$0.28 · price floor Llama 4 Maverick 1352 ███████████ Meta's last open model
Full Leaderboard
| # | Model | Quality | Arena ELO | Speed | Price | Context | Value | Released |
|---|---|---|---|---|---|---|---|---|
| 1 | Anthropic · Frontier agentic coding & knowledge work | 100 | 1525 | 58 t/s | $10 / $50 | 1M | 3.3 | Jun 2026 |
| 2 | Anthropic · Ceiling capability (limited access) | 100 | 1531 | 56 t/s | $10 / $50 | 1M | 3.3 | Jun 2026 |
| 3 | Anthropic · Frontier agentic coding & reasoning | 99 | 1522 | 74 t/s | $5 / $25 | 1M | 6.6 | Jul 2026 |
| 4 | Anthropic · Coding, agents & computer use | 98 | 1512 | 72 t/s | $5 / $25 | 1M | 6.5 | May 2026 |
| 5 | OpenAI · General frontier reasoning & tools | 98 | 1514 | 96 t/s | $5 / $30 | 1M | 5.6 | Jul 2026 |
| 6 | OpenAI · Frontier general purpose | 97 | 1506 | 70 t/s | $5 / $30 | 1M | 5.5 | Apr 2026 |
| 7 | Kimi K3OSS Moonshot AI · Open-weight coding frontier | 97 | 1500 | 55 t/s | $3 / $15 | 1M | 10.8 | Jul 2026 |
| 8 | OpenAI · Reasoning at any cost | 96 | 1510 | 68 t/s | $30 / $180 | 1M | 0.9 | Apr 2026 |
| 9 | Anthropic · Coding & agentic workflows | 96 | 1505 | 68 t/s | $5 / $25 | 1M | 6.4 | Apr 2026 |
| 10 | Google · Science & long-context | 96 | 1505 | 131 t/s | $2 / $12 | 1M | 13.7 | Apr 2026 |
| 11 | Qwen3.8 MaxOSS Alibaba Cloud · Multimodal APAC frontier (hosted); text-only open weights | 96 | 1491 | 86 t/s | $2 / $6 | 1M | 24.0 | Aug 2026 |
| 12 | OpenAI · Hard reasoning | 94 | 1370 | 68 t/s | $2 / $8 | 200K | 18.8 | Apr 2025 |
| 13 | Alibaba Cloud · Long autonomous agentic runs | 94 | 1488 | 90 t/s | $2.5 / $7.5 | 1M | 18.8 | May 2026 |
| 14 | xAI · Cheap output tokens & real-time info | 94 | 1499 | 88 t/s | $2 / $6 | 500K | 23.5 | Jul 2026 |
| 15 | xAI · Agentic tasks & real-time info | 93 | 1496 | 83 t/s | $1.25 / $2.5 | 1M | 49.6 | May 2026 |
| 16 | OpenAI · Default production tier | 93 | 1489 | 118 t/s | $2 / $12 | 1M | 13.3 | Jul 2026 |
| 17 | Anthropic · Balanced agents & coding value | 93 | 1479 | 98 t/s | $3 / $15 | 1M | 10.3 | Jun 2026 |
| 18 | Google · Multimodal + value | 92 | 1345 | 87 t/s | $1.25 / $10 | 1M | 16.4 | Mar 2025 |
| 19 | Moonshot AI · Frontier quality at low cost | 92 | 1466 | 48 t/s | $0.73 / $3.49 | 256K | 43.6 | Apr 2026 |
| 20 | Anthropic · Complex analysis | 91 | 1360 | 52 t/s | $15 / $75 | 200K | 2.0 | May 2025 |
| 21 | DeepSeek R1OSS DeepSeek · Cheap reasoning | 91 | 1350 | 35 t/s | $0.55 / $2.19 | 128K | 66.4 | Jan 2025 |
| 22 | GLM-5.2OSS Z.ai (Zhipu AI) · Fastest self-hostable frontier model | 91 | 1483 | 168 t/s | $1.4 / $4.4 | 1M | 31.4 | Jun 2026 |
| 23 | Anthropic · Coding & balance | 90 | 1467 | 73 t/s | $3 / $15 | 1M | 10.0 | Feb 2026 |
| 24 | Google · Fast multimodal at mid-tier price | 90 | 1481 | 148 t/s | $0.75 / $3.75 | 1M | 40.0 | Jul 2026 |
| 25 | OpenAI · Long context | 89 | 1310 | 120 t/s | $2 / $8 | 1M | 17.8 | Apr 2025 |
| 26 | DeepSeek · Cheap frontier-adjacent tokens (GA build) | 89 | 1461 | 62 t/s | $0.66 / $1.98 | 1M | 67.4 | Aug 2026 |
| 27 | MiniMax M3OSS MiniMax · Open-weight agentic coding | 89 | 1455 | 80 t/s | $0.6 / $2.4 | 1M | 59.3 | Jun 2026 |
| 28 | OpenAI · Reasoning & math | 88 | 1305 | 155 t/s | $1.1 / $4.4 | 200K | 32.0 | Jan 2025 |
| 29 | Anthropic · Coding & balance | 88 | 1320 | 95 t/s | $3 / $15 | 200K | 9.8 | May 2025 |
| 30 | GLM-5.1OSS Z.ai (Zhipu AI) · Open-weight agentic & tool use | 88 | 1467 | 48 t/s | $0.98 / $3.08 | 200K | 43.3 | Apr 2026 |
| 31 | xAI · Real-time info | 87 | 1330 | 82 t/s | $3 / $15 | 131K | 9.7 | Feb 2025 |
| 32 | DeepSeek V3OSS DeepSeek · Best open-source value | 86 | 1310 | 62 t/s | $0.27 / $1.1 | 128K | 125.5 | Mar 2025 |
| 33 | Alibaba Cloud · Multilingual & APAC | 86 | 1448 | 124 t/s | $1.4 / $5.6 | 256K | 24.6 | Apr 2026 |
| 34 | OpenAI · General purpose | 85 | 1285 | 109 t/s | $2.5 / $10 | 128K | 13.6 | May 2024 |
| 35 | DeepSeek · Cheapest tokens, period | 84 | 1441 | 105 t/s | $0.14 / $0.28 | 1M | 400.0 | Apr 2026 |
| 36 | OpenAI · High-volume cheap frontier-family | 83 | 1436 | 186 t/s | $0.2 / $1.2 | 1M | 118.6 | Jul 2026 |
| 37 | Meta · Open-source value | 80 | 1260 | 135 t/s | $0.2 / $0.6 | 1M | 200.0 | Apr 2025 |
| 38 | Qwen 2.5 72BOSS Alibaba Cloud · Open-source flagship | 80 | 1255 | 85 t/s | $0.3 / $0.9 | 131K | 133.3 | Sep 2024 |
| 39 | Qwen3.8 27BOSS Alibaba Cloud · Locally-runnable open-weight (~17GB VRAM at 4-bit) | 80 | — | — | Self-host | 262K | , | Aug 2026 |
| 40 | Mistral AI · Multilingual | 79 | 1250 | 78 t/s | $2 / $6 | 128K | 19.8 | Nov 2024 |
| 41 | Google · High-volume cheap multimodal | 79 | 1418 | 192 t/s | $0.3 / $2.5 | 1M | 56.4 | Jul 2026 |
| 42 | NVIDIA · Efficient single-GPU deployment, 1M context | 79 | — | — | Self-host | 1M | , | Aug 2026 |
| 43 | xAI · Budget reasoning | 78 | 1275 | 165 t/s | $0.3 / $0.5 | 131K | 195.0 | Feb 2025 |
| 44 | Perplexity · Search + citations | 78 | — | 65 t/s | $3 / $15 | 200K | 8.7 | Feb 2025 |
| 45 | Mistral AI · Code generation | 76 | — | 195 t/s | $0.3 / $0.9 | 256K | 126.7 | Jan 2025 |
| 46 | NVIDIA · Open multimodal | 76 | 1361 | 158 t/s | Self-host | 256K | , | Apr 2026 |
| 47 | Hunyuan Hy3OSS Tencent · Long-context retrieval & agentic search, no license restrictions | 76 | — | — | Self-host | 256K | , | Jul 2026 |
| 48 | Anthropic · Speed & cost | 75 | 1230 | 172 t/s | $0.8 / $4 | 200K | 31.3 | Oct 2024 |
| 49 | Gemma 4 27BOSS Google · Self-hosted general purpose | 75 | 1351 | 142 t/s | Self-host | 128K | , | Apr 2026 |
| 50 | Google · Fastest + cheapest | 74 | 1240 | 244 t/s | $0.1 / $0.4 | 1M | 296.0 | Feb 2025 |
| 51 | Alibaba Cloud · Open-source coding | 74 | — | 125 t/s | $0.15 / $0.45 | 131K | 246.7 | Nov 2024 |
| 52 | Ant Group · Efficient MoE alternative to trillion-param flagships | 74 | — | — | Self-host | 256K | , | Aug 2026 |
| 53 | OpenAI · High throughput | 72 | 1216 | 183 t/s | $0.15 / $0.6 | 128K | 192.0 | Jul 2024 |
| 54 | Meta · Longest context | 71 | 1195 | 198 t/s | $0.15 / $0.4 | 10M | 258.2 | Apr 2025 |
| 55 | Amazon · AWS ecosystem | 70 | — | 110 t/s | $0.8 / $3.2 | 300K | 35.0 | Dec 2024 |
| 56 | Cohere · Enterprise RAG | 68 | 1170 | 72 t/s | $2.5 / $10 | 128K | 10.9 | Aug 2024 |
What to do this quarter
- Update bookmarks and citations. Internal eval-spec docs and procurement RFPs that reference "lmsys.org" should be updated to lmarena.ai. The data continues at the new domain.
- Pull from the right board. Coding teams should cite the coding Arena Elo (Moonshot's Kimi K3 leads the Frontend Code Arena at ~1,679, ahead of Claude Fable 5, GPT-5.6 Sol, and Opus 4.8). Generic chat teams should cite the text leaderboard (Claude Fable 5 and Opus 5 lead at ~1522-1525, with GPT-5.6 Sol, Gemini 3.1 Pro Preview, and Kimi K3 clustered near 1500).
- Build dual-vendor capability. The top four models are within 40 Elo of each other. Treat them as interchangeable on capability and optimise for switching cost.
- Pair Arena scores with workload-specific evals. Arena rewards short-conversation polish. Long-context, tool-use, and domain-specific tasks need their own measurement.
- Track the open-weight gap. DeepSeek V4 Pro under Apache 2.0 sits at 1462 Elo, within 38 points of the text leader. The gap is the smallest it has ever been.
- Watch GPT-5.5 Pro pricing. At $30/$180 per 1M tokens, paying for the top of LMArena.ai now costs 200x more per token than the cheapest tier. The cost curve is steepening.
- Re-baseline at every model launch. Tokenizer changes between releases (as seen across the Claude Opus 4.6 → 4.7 → 4.8 line) shift effective cost without shifting list price.
Related reading
- LMArena Explained: what LMArena is, how to read Arena Elo
- AI Model Leaderboard: full quality, speed, pricing comparison
- LLM Leaderboard
- LMSys Arena Leaderboard May 2026
- LMArena Elo Explained for Enterprise Buyers
Teams running side-by-side evals against multiple LMArena.ai leaders typically expose them through Swfte Connect as a single endpoint, then run their own internal Elo on production prompts. That is the only way to verify whether public Arena rank translates to your workload.