Leaderboard

LLM Capability Leaderboard

Every major model across the benchmarks researchers cite. Sort by any column. Rows marked imported are from the public leaderboards of the benchmark maintainers; Swfte’s own runs replace them as they complete.

Updated 2026-08-04 · Methodology

Beat the benchmark

Don't just read the scores: run them

ModelProviderHuman-LikeARC-AGI-2HLEGAIASimpleBenchGPQA-DiamondMMLU-ProHuman-Like ThinkingSource
Claude Opus 4.6Anthropic::::::::imported
Claude Sonnet 4.6Anthropic::::::::imported
Claude Haiku 4.5Anthropic::::::::imported
GPT-5OpenAI::::::::imported
GPT-4.5OpenAI::::::::imported
o3-miniOpenAI::::::::imported
Gemini 2.5 ProGoogle::::::::imported
Gemini 2.5 FlashGoogle::::::::imported
Llama 4 405BMeta::::::::imported
Llama 4 70BMeta::::::::imported
Mistral Large 2Mistral::::::::imported
Mistral Small 3Mistral::::::::imported
DeepSeek V3DeepSeek::::::::imported
DeepSeek R1DeepSeek::::::::imported
Qwen 3Alibaba::::::::imported
Command R+Cohere::::::::imported
Kimi K2Moonshot::::::::imported
Grok 3xAI::::::::imported
Jamba 1.5AI21::::::::imported
Phi-4Microsoft::::::::imported
Gemma 3Google::::::::imported