Benchmarks / From signal to selection

Start with the task.
Then read the score.

A benchmark helps narrow a question. Your evaluation should connect that signal to the work the model must actually perform.

Reasoning

Check task difficulty, scoring and the model configuration. Read the capability methodology before using a composite as a shortcut.

Read the relevant reference
Published collection

The source library.

Snapshot · 2026-09-19. Read each reference for its original claims, dates and limitations.

01

Claude Opus 4.6

Anthropic · v0 docs · Updated 2026-09-19

Based on published documentation. Full audit in progress (0%).

02

Claude Sonnet 4.6

Anthropic · v0 docs · Updated 2026-09-19

Based on published documentation. Full audit in progress (0%).

03

Claude Haiku 4.5

Anthropic · v0 docs · Updated 2026-09-19

Based on published documentation. Full audit in progress (0%).

04

GPT-5

OpenAI · v0 docs · Updated 2026-09-19

Based on published documentation. Full audit in progress (0%).

05

GPT-4.5

OpenAI · v0 docs · Updated 2026-09-19

Based on published documentation. Full audit in progress (0%).

06

o3-mini

OpenAI · v0 docs · Updated 2026-09-19

Based on published documentation. Full audit in progress (0%).

07

Gemini 2.5 Pro

Google · v0 docs · Updated 2026-09-19

Based on published documentation. Full audit in progress (0%).

08

Gemini 2.5 Flash

Google · v0 docs · Updated 2026-09-19

Based on published documentation. Full audit in progress (0%).

09

Llama 4 405B

Meta · v0 docs · Updated 2026-09-19

Based on published documentation. Full audit in progress (0%).

10

Llama 4 70B

Meta · v0 docs · Updated 2026-09-19

Based on published documentation. Full audit in progress (0%).

11

Mistral Large 2

Mistral · v0 docs · Updated 2026-09-19

Based on published documentation. Full audit in progress (0%).

12

Mistral Small 3

Mistral · v0 docs · Updated 2026-09-19

Based on published documentation. Full audit in progress (0%).

13

DeepSeek V3

DeepSeek · v0 docs · Updated 2026-09-19

Based on published documentation. Full audit in progress (0%).

14

DeepSeek R1

DeepSeek · v0 docs · Updated 2026-09-19

Based on published documentation. Full audit in progress (0%).

15

Qwen 3

Alibaba · v0 docs · Updated 2026-09-19

Based on published documentation. Full audit in progress (0%).

16

Command R+

Cohere · v0 docs · Updated 2026-09-19

Based on published documentation. Full audit in progress (0%).

17

Kimi K2

Moonshot · v0 docs · Updated 2026-09-19

Based on published documentation. Full audit in progress (0%).

18

Grok 3

xAI · v0 docs · Updated 2026-09-19

Based on published documentation. Full audit in progress (0%).

19

Jamba 1.5

AI21 · v0 docs · Updated 2026-09-19

Based on published documentation. Full audit in progress (0%).

20

Phi-4

Microsoft · v0 docs · Updated 2026-09-19

Based on published documentation. Full audit in progress (0%).

21

Gemma 3

Google · v0 docs · Updated 2026-09-19

Based on published documentation. Full audit in progress (0%).

21 references

Frequently asked questions

01Which benchmarks are run?

The full adopted academic battery, ARC-AGI-2, HLE (Humanity's Last Exam), GAIA, SimpleBench, GPQA Diamond, MMLU-Pro, plus our own composites: Rationale Integrity (does the reasoning trace match the answer), Abstention (does the model refuse when it should), and the Human-Like Thinking score (aggregate across axes most predictive of agentic competence).

02How often does the leaderboard refresh?

On every model version bump and weekly otherwise. Each row shows the updatedAt timestamp of its last full benchmark pass.

03Why both academic + proprietary benchmarks?

Academic benchmarks have known training-set contamination risks: top models often hit ceiling on widely-cited tests. Our proprietary composites use unseen probes and behavioural traces that resist contamination, giving the leaderboard a longer signal-shelf-life.

04How does the Human-Like Thinking score work?

A weighted aggregate across reasoning, planning, calibrated uncertainty, abstention, and rationale-integrity axes. Tuned to correlate with downstream agentic task performance, not just multiple-choice accuracy.

05Can I download the raw results?

Yes: every benchmark row exports per-task scores plus rationale traces from the methodology page.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.