LLM Leaderboard, September 2026
Large language models ranked by LMSys Arena Elo, MMLU, HumanEval, MATH, pricing, and inference speed. Refreshed regularly with live data from official provider pricing pages, Artificial Analysis, and the Arena.
What is "the best LLM" in September 2026?
The honest answer is "depends on the workload." For chat and general reasoning, the LMSys text Arena leader rotates monthly: the September 2026 snapshot below shows the current top model and its Elo. For coding-specific work, the LMSys coding Arena has its own leader. For value (quality-per-dollar), open-weight models under permissive licenses still win by a wide margin. The race at the top is tighter than at any point since the original GPT-4 launch, and switching costs are now the buyer's biggest risk, not capability.
| # | Model | Quality | Arena ELO | Speed | Price | Context | Value | Released |
|---|---|---|---|---|---|---|---|---|
| 1 | Anthropic · Frontier agentic coding & knowledge work | 100 | 1525 | 58 t/s | $10 / $50 | 1M | 3.3 | Jun 2026 |
| 2 | Anthropic · Ceiling capability (limited access) | 100 | 1531 | 56 t/s | $10 / $50 | 1M | 3.3 | Jun 2026 |
| 3 | Anthropic · Frontier agentic coding & reasoning | 99 | 1522 | 74 t/s | $5 / $25 | 1M | 6.6 | Jul 2026 |
| 4 | Anthropic · Coding, agents & computer use | 98 | 1512 | 72 t/s | $5 / $25 | 1M | 6.5 | May 2026 |
| 5 | OpenAI · General frontier reasoning & tools | 98 | 1514 | 96 t/s | $5 / $30 | 1M | 5.6 | Jul 2026 |
| 6 | OpenAI · Frontier general purpose | 97 | 1506 | 70 t/s | $5 / $30 | 1M | 5.5 | Apr 2026 |
| 7 | Kimi K3OSS Moonshot AI · Open-weight coding frontier | 97 | 1500 | 55 t/s | $3 / $15 | 1M | 10.8 | Jul 2026 |
| 8 | OpenAI · Reasoning at any cost | 96 | 1510 | 68 t/s | $30 / $180 | 1M | 0.9 | Apr 2026 |
| 9 | Anthropic · Coding & agentic workflows | 96 | 1505 | 68 t/s | $5 / $25 | 1M | 6.4 | Apr 2026 |
| 10 | Google · Science & long-context | 96 | 1505 | 131 t/s | $2 / $12 | 1M | 13.7 | Apr 2026 |
| 11 | Qwen3.8 MaxOSS Alibaba Cloud · Multimodal APAC frontier (hosted); text-only open weights | 96 | 1491 | 86 t/s | $2 / $6 | 1M | 24.0 | Aug 2026 |
| 12 | OpenAI · Hard reasoning | 94 | 1370 | 68 t/s | $2 / $8 | 200K | 18.8 | Apr 2025 |
| 13 | Alibaba Cloud · Long autonomous agentic runs | 94 | 1488 | 90 t/s | $2.5 / $7.5 | 1M | 18.8 | May 2026 |
| 14 | xAI · Cheap output tokens & real-time info | 94 | 1499 | 88 t/s | $2 / $6 | 500K | 23.5 | Jul 2026 |
| 15 | xAI · Agentic tasks & real-time info | 93 | 1496 | 83 t/s | $1.25 / $2.5 | 1M | 49.6 | May 2026 |
| 16 | OpenAI · Default production tier | 93 | 1489 | 118 t/s | $2 / $12 | 1M | 13.3 | Jul 2026 |
| 17 | Anthropic · Balanced agents & coding value | 93 | 1479 | 98 t/s | $3 / $15 | 1M | 10.3 | Jun 2026 |
| 18 | Google · Multimodal + value | 92 | 1345 | 87 t/s | $1.25 / $10 | 1M | 16.4 | Mar 2025 |
| 19 | Moonshot AI · Frontier quality at low cost | 92 | 1466 | 48 t/s | $0.73 / $3.49 | 256K | 43.6 | Apr 2026 |
| 20 | Anthropic · Complex analysis | 91 | 1360 | 52 t/s | $15 / $75 | 200K | 2.0 | May 2025 |
| 21 | DeepSeek R1OSS DeepSeek · Cheap reasoning | 91 | 1350 | 35 t/s | $0.55 / $2.19 | 128K | 66.4 | Jan 2025 |
| 22 | GLM-5.2OSS Z.ai (Zhipu AI) · Fastest self-hostable frontier model | 91 | 1483 | 168 t/s | $1.4 / $4.4 | 1M | 31.4 | Jun 2026 |
| 23 | Anthropic · Coding & balance | 90 | 1467 | 73 t/s | $3 / $15 | 1M | 10.0 | Feb 2026 |
| 24 | Google · Fast multimodal at mid-tier price | 90 | 1481 | 148 t/s | $0.75 / $3.75 | 1M | 40.0 | Jul 2026 |
| 25 | OpenAI · Long context | 89 | 1310 | 120 t/s | $2 / $8 | 1M | 17.8 | Apr 2025 |
| 26 | DeepSeek · Cheap frontier-adjacent tokens (GA build) | 89 | 1461 | 62 t/s | $0.66 / $1.98 | 1M | 67.4 | Aug 2026 |
| 27 | MiniMax M3OSS MiniMax · Open-weight agentic coding | 89 | 1455 | 80 t/s | $0.6 / $2.4 | 1M | 59.3 | Jun 2026 |
| 28 | OpenAI · Reasoning & math | 88 | 1305 | 155 t/s | $1.1 / $4.4 | 200K | 32.0 | Jan 2025 |
| 29 | Anthropic · Coding & balance | 88 | 1320 | 95 t/s | $3 / $15 | 200K | 9.8 | May 2025 |
| 30 | GLM-5.1OSS Z.ai (Zhipu AI) · Open-weight agentic & tool use | 88 | 1467 | 48 t/s | $0.98 / $3.08 | 200K | 43.3 | Apr 2026 |
| 31 | xAI · Real-time info | 87 | 1330 | 82 t/s | $3 / $15 | 131K | 9.7 | Feb 2025 |
| 32 | DeepSeek V3OSS DeepSeek · Best open-source value | 86 | 1310 | 62 t/s | $0.27 / $1.1 | 128K | 125.5 | Mar 2025 |
| 33 | Alibaba Cloud · Multilingual & APAC | 86 | 1448 | 124 t/s | $1.4 / $5.6 | 256K | 24.6 | Apr 2026 |
| 34 | OpenAI · General purpose | 85 | 1285 | 109 t/s | $2.5 / $10 | 128K | 13.6 | May 2024 |
| 35 | DeepSeek · Cheapest tokens, period | 84 | 1441 | 105 t/s | $0.14 / $0.28 | 1M | 400.0 | Apr 2026 |
| 36 | OpenAI · High-volume cheap frontier-family | 83 | 1436 | 186 t/s | $0.2 / $1.2 | 1M | 118.6 | Jul 2026 |
| 37 | Meta · Open-source value | 80 | 1260 | 135 t/s | $0.2 / $0.6 | 1M | 200.0 | Apr 2025 |
| 38 | Qwen 2.5 72BOSS Alibaba Cloud · Open-source flagship | 80 | 1255 | 85 t/s | $0.3 / $0.9 | 131K | 133.3 | Sep 2024 |
| 39 | Qwen3.8 27BOSS Alibaba Cloud · Locally-runnable open-weight (~17GB VRAM at 4-bit) | 80 | — | — | Self-host | 262K | , | Aug 2026 |
| 40 | Mistral AI · Multilingual | 79 | 1250 | 78 t/s | $2 / $6 | 128K | 19.8 | Nov 2024 |
| 41 | Google · High-volume cheap multimodal | 79 | 1418 | 192 t/s | $0.3 / $2.5 | 1M | 56.4 | Jul 2026 |
| 42 | NVIDIA · Efficient single-GPU deployment, 1M context | 79 | — | — | Self-host | 1M | , | Aug 2026 |
| 43 | xAI · Budget reasoning | 78 | 1275 | 165 t/s | $0.3 / $0.5 | 131K | 195.0 | Feb 2025 |
| 44 | Perplexity · Search + citations | 78 | — | 65 t/s | $3 / $15 | 200K | 8.7 | Feb 2025 |
| 45 | Mistral AI · Code generation | 76 | — | 195 t/s | $0.3 / $0.9 | 256K | 126.7 | Jan 2025 |
| 46 | NVIDIA · Open multimodal | 76 | 1361 | 158 t/s | Self-host | 256K | , | Apr 2026 |
| 47 | Hunyuan Hy3OSS Tencent · Long-context retrieval & agentic search, no license restrictions | 76 | — | — | Self-host | 256K | , | Jul 2026 |
| 48 | Anthropic · Speed & cost | 75 | 1230 | 172 t/s | $0.8 / $4 | 200K | 31.3 | Oct 2024 |
| 49 | Gemma 4 27BOSS Google · Self-hosted general purpose | 75 | 1351 | 142 t/s | Self-host | 128K | , | Apr 2026 |
| 50 | Google · Fastest + cheapest | 74 | 1240 | 244 t/s | $0.1 / $0.4 | 1M | 296.0 | Feb 2025 |
| 51 | Alibaba Cloud · Open-source coding | 74 | — | 125 t/s | $0.15 / $0.45 | 131K | 246.7 | Nov 2024 |
| 52 | Ant Group · Efficient MoE alternative to trillion-param flagships | 74 | — | — | Self-host | 256K | , | Aug 2026 |
| 53 | OpenAI · High throughput | 72 | 1216 | 183 t/s | $0.15 / $0.6 | 128K | 192.0 | Jul 2024 |
| 54 | Meta · Longest context | 71 | 1195 | 198 t/s | $0.15 / $0.4 | 10M | 258.2 | Apr 2025 |
| 55 | Amazon · AWS ecosystem | 70 | — | 110 t/s | $0.8 / $3.2 | 300K | 35.0 | Dec 2024 |
| 56 | Cohere · Enterprise RAG | 68 | 1170 | 72 t/s | $2.5 / $10 | 128K | 10.9 | Aug 2024 |
How the LLM leaderboard works
We pull official provider pricing every 24 hours, Artificial Analysis benchmark snapshots weekly, and LMSys Arena Elo as it publishes. The composite quality index is a 0-100 normalization over MMLU Pro, HumanEval, and MATH, weighted by recency and cross-validated against Arena Elo. We do not accept vendor-supplied numbers without an independent reference.
Where the leaderboard is wrong
No leaderboard predicts your production accuracy. LMSys Arena rewards style and short-conversation polish; a top-Arena model can still under-perform on your specific function-calling schema or long-context retrieval workload. Build an internal eval harness before you commit. See our LMArena Elo explained and LLM routing writeups for the deep-dive.
Related rankings
- AI Model Leaderboard: same data, broader entry point
- Models Leaderboard
- GenAI Leaderboard
- AI Vendor Lock-in Leaderboard