AI Model Leaderboard, September 2026
Every major AI model ranked by quality, speed, pricing, and value. Filter by category, sort by any metric, and find the right model for your use case. Live data refreshed regularly with LMSys Arena Elo, official provider pricing, and Artificial Analysis benchmarks.
In short:As of September 2026, Claude Fable 5 leads the AI model leaderboard at 100/100 on our composite quality index, across 56+ models ranked by quality, price, speed, and value. The best pick depends on the workload: sort the table by the metric that matters to you.
Claude Fable 5
100/100
Claude Mythos 5
100/100
Claude Opus 5
99/100
Stop reading: start ranking
Three ways to put this leaderboard to work. Pick any one, they all start with a free Swfte account, no card required.
September 2026: Top Models, Best Value, Fastest Inference
The September 2026 ranking covers 56 models across LMSys Arena Elo, MMLU Pro, HumanEval, MATH, pricing, and inference speed. Top of the table: Claude Fable 5 at 100/100 quality. The full table below is sortable by any metric. Live data is refreshed regularly from official provider pricing pages and the public Arena.
Top 5 by Quality Index
- Claude Fable 5: 100/100
- Claude Mythos 5: 100/100
- Claude Opus 5: 99/100
- Claude Opus 4.8: 98/100
- GPT-5.6 Sol: 98/100
Best Price-to-Quality
- DeepSeek V4 Flash: $0.28/1M out
- Gemini 2.0 Flash: $0.4/1M out
- Llama 4 Scout: $0.4/1M out
- Qwen 2.5 Coder 32B: $0.45/1M out
- Grok 3 Mini: $0.5/1M out
See our LMSys Arena deep dive and the monthly release roundup.
| # | Model | Quality | Arena ELO | Speed | Price | Context | Value | Released |
|---|---|---|---|---|---|---|---|---|
| 1 | Anthropic · Frontier agentic coding & knowledge work | 100 | 1525 | 58 t/s | $10 / $50 | 1M | 3.3 | Jun 2026 |
| 2 | Anthropic · Ceiling capability (limited access) | 100 | 1531 | 56 t/s | $10 / $50 | 1M | 3.3 | Jun 2026 |
| 3 | Anthropic · Frontier agentic coding & reasoning | 99 | 1522 | 74 t/s | $5 / $25 | 1M | 6.6 | Jul 2026 |
| 4 | Anthropic · Coding, agents & computer use | 98 | 1512 | 72 t/s | $5 / $25 | 1M | 6.5 | May 2026 |
| 5 | OpenAI · General frontier reasoning & tools | 98 | 1514 | 96 t/s | $5 / $30 | 1M | 5.6 | Jul 2026 |
| 6 | OpenAI · Frontier general purpose | 97 | 1506 | 70 t/s | $5 / $30 | 1M | 5.5 | Apr 2026 |
| 7 | Kimi K3OSS Moonshot AI · Open-weight coding frontier | 97 | 1500 | 55 t/s | $3 / $15 | 1M | 10.8 | Jul 2026 |
| 8 | OpenAI · Reasoning at any cost | 96 | 1510 | 68 t/s | $30 / $180 | 1M | 0.9 | Apr 2026 |
| 9 | Anthropic · Coding & agentic workflows | 96 | 1505 | 68 t/s | $5 / $25 | 1M | 6.4 | Apr 2026 |
| 10 | Google · Science & long-context | 96 | 1505 | 131 t/s | $2 / $12 | 1M | 13.7 | Apr 2026 |
| 11 | Qwen3.8 MaxOSS Alibaba Cloud · Multimodal APAC frontier (hosted); text-only open weights | 96 | 1491 | 86 t/s | $2 / $6 | 1M | 24.0 | Aug 2026 |
| 12 | OpenAI · Hard reasoning | 94 | 1370 | 68 t/s | $2 / $8 | 200K | 18.8 | Apr 2025 |
| 13 | Alibaba Cloud · Long autonomous agentic runs | 94 | 1488 | 90 t/s | $2.5 / $7.5 | 1M | 18.8 | May 2026 |
| 14 | xAI · Cheap output tokens & real-time info | 94 | 1499 | 88 t/s | $2 / $6 | 500K | 23.5 | Jul 2026 |
| 15 | xAI · Agentic tasks & real-time info | 93 | 1496 | 83 t/s | $1.25 / $2.5 | 1M | 49.6 | May 2026 |
| 16 | OpenAI · Default production tier | 93 | 1489 | 118 t/s | $2 / $12 | 1M | 13.3 | Jul 2026 |
| 17 | Anthropic · Balanced agents & coding value | 93 | 1479 | 98 t/s | $3 / $15 | 1M | 10.3 | Jun 2026 |
| 18 | Google · Multimodal + value | 92 | 1345 | 87 t/s | $1.25 / $10 | 1M | 16.4 | Mar 2025 |
| 19 | Moonshot AI · Frontier quality at low cost | 92 | 1466 | 48 t/s | $0.73 / $3.49 | 256K | 43.6 | Apr 2026 |
| 20 | Anthropic · Complex analysis | 91 | 1360 | 52 t/s | $15 / $75 | 200K | 2.0 | May 2025 |
| 21 | DeepSeek R1OSS DeepSeek · Cheap reasoning | 91 | 1350 | 35 t/s | $0.55 / $2.19 | 128K | 66.4 | Jan 2025 |
| 22 | GLM-5.2OSS Z.ai (Zhipu AI) · Fastest self-hostable frontier model | 91 | 1483 | 168 t/s | $1.4 / $4.4 | 1M | 31.4 | Jun 2026 |
| 23 | Anthropic · Coding & balance | 90 | 1467 | 73 t/s | $3 / $15 | 1M | 10.0 | Feb 2026 |
| 24 | Google · Fast multimodal at mid-tier price | 90 | 1481 | 148 t/s | $0.75 / $3.75 | 1M | 40.0 | Jul 2026 |
| 25 | OpenAI · Long context | 89 | 1310 | 120 t/s | $2 / $8 | 1M | 17.8 | Apr 2025 |
| 26 | DeepSeek · Cheap frontier-adjacent tokens (GA build) | 89 | 1461 | 62 t/s | $0.66 / $1.98 | 1M | 67.4 | Aug 2026 |
| 27 | MiniMax M3OSS MiniMax · Open-weight agentic coding | 89 | 1455 | 80 t/s | $0.6 / $2.4 | 1M | 59.3 | Jun 2026 |
| 28 | OpenAI · Reasoning & math | 88 | 1305 | 155 t/s | $1.1 / $4.4 | 200K | 32.0 | Jan 2025 |
| 29 | Anthropic · Coding & balance | 88 | 1320 | 95 t/s | $3 / $15 | 200K | 9.8 | May 2025 |
| 30 | GLM-5.1OSS Z.ai (Zhipu AI) · Open-weight agentic & tool use | 88 | 1467 | 48 t/s | $0.98 / $3.08 | 200K | 43.3 | Apr 2026 |
| 31 | xAI · Real-time info | 87 | 1330 | 82 t/s | $3 / $15 | 131K | 9.7 | Feb 2025 |
| 32 | DeepSeek V3OSS DeepSeek · Best open-source value | 86 | 1310 | 62 t/s | $0.27 / $1.1 | 128K | 125.5 | Mar 2025 |
| 33 | Alibaba Cloud · Multilingual & APAC | 86 | 1448 | 124 t/s | $1.4 / $5.6 | 256K | 24.6 | Apr 2026 |
| 34 | OpenAI · General purpose | 85 | 1285 | 109 t/s | $2.5 / $10 | 128K | 13.6 | May 2024 |
| 35 | DeepSeek · Cheapest tokens, period | 84 | 1441 | 105 t/s | $0.14 / $0.28 | 1M | 400.0 | Apr 2026 |
| 36 | OpenAI · High-volume cheap frontier-family | 83 | 1436 | 186 t/s | $0.2 / $1.2 | 1M | 118.6 | Jul 2026 |
| 37 | Meta · Open-source value | 80 | 1260 | 135 t/s | $0.2 / $0.6 | 1M | 200.0 | Apr 2025 |
| 38 | Qwen 2.5 72BOSS Alibaba Cloud · Open-source flagship | 80 | 1255 | 85 t/s | $0.3 / $0.9 | 131K | 133.3 | Sep 2024 |
| 39 | Qwen3.8 27BOSS Alibaba Cloud · Locally-runnable open-weight (~17GB VRAM at 4-bit) | 80 | — | — | Self-host | 262K | , | Aug 2026 |
| 40 | Mistral AI · Multilingual | 79 | 1250 | 78 t/s | $2 / $6 | 128K | 19.8 | Nov 2024 |
| 41 | Google · High-volume cheap multimodal | 79 | 1418 | 192 t/s | $0.3 / $2.5 | 1M | 56.4 | Jul 2026 |
| 42 | NVIDIA · Efficient single-GPU deployment, 1M context | 79 | — | — | Self-host | 1M | , | Aug 2026 |
| 43 | xAI · Budget reasoning | 78 | 1275 | 165 t/s | $0.3 / $0.5 | 131K | 195.0 | Feb 2025 |
| 44 | Perplexity · Search + citations | 78 | — | 65 t/s | $3 / $15 | 200K | 8.7 | Feb 2025 |
| 45 | Mistral AI · Code generation | 76 | — | 195 t/s | $0.3 / $0.9 | 256K | 126.7 | Jan 2025 |
| 46 | NVIDIA · Open multimodal | 76 | 1361 | 158 t/s | Self-host | 256K | , | Apr 2026 |
| 47 | Hunyuan Hy3OSS Tencent · Long-context retrieval & agentic search, no license restrictions | 76 | — | — | Self-host | 256K | , | Jul 2026 |
| 48 | Anthropic · Speed & cost | 75 | 1230 | 172 t/s | $0.8 / $4 | 200K | 31.3 | Oct 2024 |
| 49 | Gemma 4 27BOSS Google · Self-hosted general purpose | 75 | 1351 | 142 t/s | Self-host | 128K | , | Apr 2026 |
| 50 | Google · Fastest + cheapest | 74 | 1240 | 244 t/s | $0.1 / $0.4 | 1M | 296.0 | Feb 2025 |
| 51 | Alibaba Cloud · Open-source coding | 74 | — | 125 t/s | $0.15 / $0.45 | 131K | 246.7 | Nov 2024 |
| 52 | Ant Group · Efficient MoE alternative to trillion-param flagships | 74 | — | — | Self-host | 256K | , | Aug 2026 |
| 53 | OpenAI · High throughput | 72 | 1216 | 183 t/s | $0.15 / $0.6 | 128K | 192.0 | Jul 2024 |
| 54 | Meta · Longest context | 71 | 1195 | 198 t/s | $0.15 / $0.4 | 10M | 258.2 | Apr 2025 |
| 55 | Amazon · AWS ecosystem | 70 | — | 110 t/s | $0.8 / $3.2 | 300K | 35.0 | Dec 2024 |
| 56 | Cohere · Enterprise RAG | 68 | 1170 | 72 t/s | $2.5 / $10 | 128K | 10.9 | Aug 2024 |
LLM Leaderboard September 2026
Large language models ranked by LMSys Arena Elo, MMLU, HumanEval, MATH, pricing, and tokens-per-second. Text-only view.
LM Leaderboard September 2026
Language model rankings: LMArena Elo, price-to-Elo ratio, and open-weight vs closed-source comparison.
LMSys Arena Leaderboard September 2026
LMArena (formerly LMSys Chatbot Arena) tracker: pairwise human preference Elo scores, refreshed as the public arena publishes.
Image Model Leaderboard 2026
Generative AI image and video models: Imagen 4, Flux 2, DALL-E 4, Stable Diffusion 4 Ultra, Sora 2 ranked by quality and cost.
Coding Model Leaderboard 2026
AI coding assistants ranked: Claude Opus, GPT-5.5, Gemini 3.1 Pro, DeepSeek V4, plus HumanEval and SWE-Bench scores.
Vendor Lock-in Leaderboard 2026
AI vendors ranked by portability: license, weight availability, fine-tuning openness, and exit cost score.
How We Rank AI Models
Our leaderboard uses a composite quality index that combines three key benchmarks: MMLU Pro (measuring knowledge and reasoning across 57 subjects), HumanEval (measuring code generation ability), and MATH (measuring mathematical problem-solving). Scores are normalized to a 0-100 scale and cross-referenced against LMSYS Chatbot Arena ELO ratings for real-world validation.
We track speed (tokens per second), time-to-first-token (TTFT), pricing, and context window size to give you a complete picture. The Value Score divides quality by cost, showing you which models deliver the most capability per dollar.
Key Trends in AI Model Performance
- A new frontier #1: Anthropic's Claude Opus 5 (24 July 2026) tops the board — a step change over Opus 4.8 on deep reasoning and long-horizon agentic work, at unchanged $5/$25 pricing and roughly half the cost of Claude Fable 5. Thinking is on by default and the prompt-cache minimum halves to 512 tokens
- OpenAI split GPT-5.6 into three tiers: Sol ($5/$30), Terra ($2/$12), and Luna ($0.20/$1.20), GA 9 July. On 30 July OpenAI cut Terra 20% and Luna 80%, citing inference work that reduced serving cost 20% — the steepest cut of the year from a US lab, and a direct answer to the Chinese open-weight tier on price
- Open weights have caught the frontier: Moonshot's Kimi K3 (2.8T MoE) ranks #3 overall on the Artificial Analysis Intelligence Index — ahead of every proprietary model except Claude Fable 5 and GPT-5.6 Sol — and took #1 on the Frontend Code Arena. It is not a lone outlier: DeepSeek V4 Pro posts 80.6% on SWE-bench Verified (matching Gemini 3.1 Pro) and MiniMax M3 80.5%, while GLM-5.2 scores 62.1% on SWE-bench Pro against GPT-5.5's 58.6% — an MIT-licensed model you can self-host outscoring a $5/$30 US flagship on agentic coding
- "Open" now means datacentre-scale: the leading open models are trillion-parameter MoEs — GLM-5.2 needs ~1TB VRAM in BF16 (~8x H200 at FP8), Kimi K3 needs 64+ accelerators. The ecosystem has bifurcated into giant open MoEs and genuinely small models (Qwen 3.6 27B, Gemma 4, Nemotron 3 Nano Omni) that run on one GPU or a phone, with little in between
- Meta left the open frontier: Llama 5 has not shipped and is now forecast for 2027. Meta pivoted to its first closed frontier model, Muse Spark (April 2026), with no weights and no architecture paper — leaving Chinese labs (DeepSeek, Moonshot, Z.ai, Alibaba, MiniMax) as the effective owners of the open-weight frontier
- Reasoning is built in: Frontier models like GPT-5.6 Sol, Claude Opus 5, and Gemini 3.1 Pro ship extended thinking by default, while deep-research variants trade latency for accuracy on the hardest tasks
- Million-token context is standard: 1M-token windows are now table stakes for flagship models across OpenAI, Anthropic, Google, and xAI
- DeepSeek is the price floor: V4 Pro sits at $0.435/$0.87 per 1M tokens — launched at $1.74/$3.48, cut 75% as a promotion, then made permanent on 31 May 2026 — and V4 Flash at $0.14/$0.28 is the cheapest usable model anywhere, with cache hits at $0.0028. Measured cost-per-task is about $0.04 on V4 Pro against $0.94 on Kimi K3 and far more on the closed frontier. One dollar buys ~1.15M output tokens on V4 Pro, ~227K on GLM-5.2, ~67K on Kimi K3. Watch one thing: DeepSeek has announced 2x peak-hour pricing for two daily windows, not yet active as of 4 August
Choosing the Right Model
There is no single "best" model: it depends on your use case. For most applications, a model routing approach works best: route simple queries to fast, cheap models and complex queries to frontier models. This gives you the best of both worlds: low cost and high quality.