Start with the task.
Then read the score.
A benchmark helps narrow a question. Your evaluation should connect that signal to the work the model must actually perform.
Reasoning
Check task difficulty, scoring and the model configuration. Read the capability methodology before using a composite as a shortcut.
Read the relevant referenceThe source library.
Snapshot · 2026-09-18. Read each reference for its original claims, dates and limitations.
Claude Opus 4.6
Anthropic · v0 docs · Updated 2026-09-18
Based on published documentation. Full audit in progress (0%).
Claude Sonnet 4.6
Anthropic · v0 docs · Updated 2026-09-18
Based on published documentation. Full audit in progress (0%).
Claude Haiku 4.5
Anthropic · v0 docs · Updated 2026-09-18
Based on published documentation. Full audit in progress (0%).
GPT-5
OpenAI · v0 docs · Updated 2026-09-18
Based on published documentation. Full audit in progress (0%).
GPT-4.5
OpenAI · v0 docs · Updated 2026-09-18
Based on published documentation. Full audit in progress (0%).
o3-mini
OpenAI · v0 docs · Updated 2026-09-18
Based on published documentation. Full audit in progress (0%).
Gemini 2.5 Pro
Google · v0 docs · Updated 2026-09-18
Based on published documentation. Full audit in progress (0%).
Gemini 2.5 Flash
Google · v0 docs · Updated 2026-09-18
Based on published documentation. Full audit in progress (0%).
Llama 4 405B
Meta · v0 docs · Updated 2026-09-18
Based on published documentation. Full audit in progress (0%).
Llama 4 70B
Meta · v0 docs · Updated 2026-09-18
Based on published documentation. Full audit in progress (0%).
Mistral Large 2
Mistral · v0 docs · Updated 2026-09-18
Based on published documentation. Full audit in progress (0%).
Mistral Small 3
Mistral · v0 docs · Updated 2026-09-18
Based on published documentation. Full audit in progress (0%).
DeepSeek V3
DeepSeek · v0 docs · Updated 2026-09-18
Based on published documentation. Full audit in progress (0%).
DeepSeek R1
DeepSeek · v0 docs · Updated 2026-09-18
Based on published documentation. Full audit in progress (0%).
Qwen 3
Alibaba · v0 docs · Updated 2026-09-18
Based on published documentation. Full audit in progress (0%).
Command R+
Cohere · v0 docs · Updated 2026-09-18
Based on published documentation. Full audit in progress (0%).
Kimi K2
Moonshot · v0 docs · Updated 2026-09-18
Based on published documentation. Full audit in progress (0%).
Grok 3
xAI · v0 docs · Updated 2026-09-18
Based on published documentation. Full audit in progress (0%).
Jamba 1.5
AI21 · v0 docs · Updated 2026-09-18
Based on published documentation. Full audit in progress (0%).
Phi-4
Microsoft · v0 docs · Updated 2026-09-18
Based on published documentation. Full audit in progress (0%).
Gemma 3
Google · v0 docs · Updated 2026-09-18
Based on published documentation. Full audit in progress (0%).
No models match your filters.
21 references
Frequently asked questions
01Which benchmarks are run?
The full adopted academic battery, ARC-AGI-2, HLE (Humanity's Last Exam), GAIA, SimpleBench, GPQA Diamond, MMLU-Pro, plus our own composites: Rationale Integrity (does the reasoning trace match the answer), Abstention (does the model refuse when it should), and the Human-Like Thinking score (aggregate across axes most predictive of agentic competence).
02How often does the leaderboard refresh?
On every model version bump and weekly otherwise. Each row shows the updatedAt timestamp of its last full benchmark pass.
03Why both academic + proprietary benchmarks?
Academic benchmarks have known training-set contamination risks: top models often hit ceiling on widely-cited tests. Our proprietary composites use unseen probes and behavioural traces that resist contamination, giving the leaderboard a longer signal-shelf-life.
04How does the Human-Like Thinking score work?
A weighted aggregate across reasoning, planning, calibrated uncertainty, abstention, and rationale-integrity axes. Tuned to correlate with downstream agentic task performance, not just multiple-choice accuracy.
05Can I download the raw results?
Yes: every benchmark row exports per-task scores plus rationale traces from the methodology page.