Ethics & trustworthiness

Testing what the labs
actually claim their models do.

Trustworthiness scorecards running live probes: the 8 DecodingTrust axes, our own Break-Free sandbox-escape harness, a Claim-Validation pass that verifies published safety assertions, and a Pressure-Drift curve showing how alignment degrades under sustained pressure.

Read the audit methodology
Three lenses for a closer look

Put the question
before the claim.

THE QUESTION / BOUNDARIES

What happens when a boundary is challenged?

Read how the published Break-Free harness probes sandbox escape, and which observations the test records.

Explore the Break-Free method
Published collection

The source library.

Snapshot · 2026-09-19. Read each scorecard for its cited evidence, dates and limitations.

01

Claude Opus 4.6 ethics audit

Anthropic · Based on published documentation (0%) · Updated 2026-09-19

02

Claude Sonnet 4.6 ethics audit

Anthropic · Based on published documentation (0%) · Updated 2026-09-19

03

Claude Haiku 4.5 ethics audit

Anthropic · Based on published documentation (0%) · Updated 2026-09-19

04

GPT-5 ethics audit

OpenAI · Based on published documentation (0%) · Updated 2026-09-19

05

GPT-4.5 ethics audit

OpenAI · Based on published documentation (0%) · Updated 2026-09-19

06

o3-mini ethics audit

OpenAI · Based on published documentation (0%) · Updated 2026-09-19

07

Gemini 2.5 Pro ethics audit

Google · Based on published documentation (0%) · Updated 2026-09-19

08

Gemini 2.5 Flash ethics audit

Google · Based on published documentation (0%) · Updated 2026-09-19

09

Llama 4 405B ethics audit

Meta · Based on published documentation (0%) · Updated 2026-09-19

10

Llama 4 70B ethics audit

Meta · Based on published documentation (0%) · Updated 2026-09-19

11

Mistral Large 2 ethics audit

Mistral · Based on published documentation (0%) · Updated 2026-09-19

12

Mistral Small 3 ethics audit

Mistral · Based on published documentation (0%) · Updated 2026-09-19

13

DeepSeek V3 ethics audit

DeepSeek · Based on published documentation (0%) · Updated 2026-09-19

14

DeepSeek R1 ethics audit

DeepSeek · Based on published documentation (0%) · Updated 2026-09-19

15

Qwen 3 ethics audit

Alibaba · Based on published documentation (0%) · Updated 2026-09-19

16

Command R+ ethics audit

Cohere · Based on published documentation (0%) · Updated 2026-09-19

17

Kimi K2 ethics audit

Moonshot · Based on published documentation (0%) · Updated 2026-09-19

18

Grok 3 ethics audit

xAI · Based on published documentation (0%) · Updated 2026-09-19

19

Jamba 1.5 ethics audit

AI21 · Based on published documentation (0%) · Updated 2026-09-19

20

Phi-4 ethics audit

Microsoft · Based on published documentation (0%) · Updated 2026-09-19

21

Gemma 3 ethics audit

Google · Based on published documentation (0%) · Updated 2026-09-19

21 references

Common questions

What the audit answers.

What does an LLM ethics audit cover?

Four pillars. (1) DecodingTrust: 8 axes of trustworthiness including toxicity, stereotype bias, adversarial robustness, out-of-distribution behaviour, privacy, machine ethics, and fairness. (2) Sandbox-escape: a live "Break-Free" harness that probes whether the model attempts to circumvent stated constraints. (3) Claim-validation: verifying every published safety assertion from the provider. (4) Pressure-drift: how alignment degrades under sustained adversarial pressure.

Why test claim validation specifically?

Model providers publish safety claims that are frequently overstated or untested. Claim-validation runs the claim against the model and reports whether the behaviour matches the stated specification: a much stronger signal than the marketing.

How does Pressure-Drift work?

A scripted adversary runs sustained conversational pressure (jailbreak attempts, social engineering, persistence). The curve shows how the model's refusal rate decays over turn count. Models that hold steady earn high marks; models that capitulate quickly are flagged.

Are these tests open?

The methodology is public. The exact probe prompts are partially gated: fully open prompts get trained against, eroding the signal. Methodology details and partial probe examples are on the methodology page.

How are ethics scores different from benchmarks?

Capability benchmarks measure what a model can do. Ethics audits measure what a model should refuse or constrain. A model can ace ARC-AGI and still fail the ethics suite. They are orthogonal axes.

See what your agents are actually doing

Nexus gives you governance, observability and spend control across every agent you run.