Governance / Decision record

Every major LLM, scored against
NIST AI RMF and the EU AI Act.

Compliance scorecards for procurement and risk teams. Each model is audited against the four NIST AI RMF functions and EU AI Act Articles 10-15, with cited evidence from published documentation and live behavioural probes.

Explore the review framework
01Purpose & scope

What decision will this assessment inform?

Name the model, intended use and the people affected. Keep a model evaluation separate from approval of a specific deployment.

02Evidence & method

Can someone else follow the assessment?

Record the methodology, source date and supporting material. Mark missing evidence explicitly.

03Ownership & review

Who owns the next decision?

Identify the reviewer, the unresolved questions and the conditions that would trigger reassessment.

Published collection

The source library.

Snapshot · 2026-09-18. Read each scorecard for its cited evidence, dates and limitations.

01

Claude Opus 4.6 governance audit

Anthropic · Based on published documentation (0%) · Updated 2026-09-18

02

Claude Sonnet 4.6 governance audit

Anthropic · Based on published documentation (0%) · Updated 2026-09-18

03

Claude Haiku 4.5 governance audit

Anthropic · Based on published documentation (0%) · Updated 2026-09-18

04

GPT-5 governance audit

OpenAI · Based on published documentation (0%) · Updated 2026-09-18

05

GPT-4.5 governance audit

OpenAI · Based on published documentation (0%) · Updated 2026-09-18

06

o3-mini governance audit

OpenAI · Based on published documentation (0%) · Updated 2026-09-18

07

Gemini 2.5 Pro governance audit

Google · Based on published documentation (0%) · Updated 2026-09-18

08

Gemini 2.5 Flash governance audit

Google · Based on published documentation (0%) · Updated 2026-09-18

09

Llama 4 405B governance audit

Meta · Based on published documentation (0%) · Updated 2026-09-18

10

Llama 4 70B governance audit

Meta · Based on published documentation (0%) · Updated 2026-09-18

11

Mistral Large 2 governance audit

Mistral · Based on published documentation (0%) · Updated 2026-09-18

12

Mistral Small 3 governance audit

Mistral · Based on published documentation (0%) · Updated 2026-09-18

13

DeepSeek V3 governance audit

DeepSeek · Based on published documentation (0%) · Updated 2026-09-18

14

DeepSeek R1 governance audit

DeepSeek · Based on published documentation (0%) · Updated 2026-09-18

15

Qwen 3 governance audit

Alibaba · Based on published documentation (0%) · Updated 2026-09-18

16

Command R+ governance audit

Cohere · Based on published documentation (0%) · Updated 2026-09-18

17

Kimi K2 governance audit

Moonshot · Based on published documentation (0%) · Updated 2026-09-18

18

Grok 3 governance audit

xAI · Based on published documentation (0%) · Updated 2026-09-18

19

Jamba 1.5 governance audit

AI21 · Based on published documentation (0%) · Updated 2026-09-18

20

Phi-4 governance audit

Microsoft · Based on published documentation (0%) · Updated 2026-09-18

21

Gemma 3 governance audit

Google · Based on published documentation (0%) · Updated 2026-09-18

21 references

Common questions

What the audit answers.

What is an LLM governance audit?

A governance audit scores a large language model against a formal framework (NIST AI RMF, the EU AI Act, ISO/IEC 42001, or sector-specific rules) and produces a procurement-grade scorecard. The output covers the four NIST functions (Govern, Map, Measure, Manage) and EU AI Act Articles 10-15: data governance, technical documentation, transparency, human oversight, accuracy, robustness, and cybersecurity.

Which models are scored?

Every major frontier LLM available via API or open weights: Claude Opus 4.7 and Sonnet 4 (Anthropic), GPT-5.5 Pro and GPT-5.5 (OpenAI), Gemini 3.1 Pro and 3.0 (Google), DeepSeek V4 Pro and V4 (DeepSeek), Llama 4 (Meta), Grok 4 (xAI), Qwen 3 (Alibaba), Mistral Large, Kimi K2.5, and Command R+. New releases are added within 30 days of GA.

How are scores calculated?

Each model is scored on cited evidence from published documentation (model cards, system cards, transparency reports) plus live behavioural probes against the audit harness. Completeness reflects how much of the NIST + EU AI Act control set has verifiable evidence; the score is a weighted aggregate across the controls.

How often are the audits refreshed?

Audits refresh on every model version bump and quarterly otherwise. Each scorecard shows updatedAt: the timestamp of the last evidence pass.

Can I use these scorecards in a procurement RFP?

Yes. That is the primary use case. Each scorecard exports as a procurement-grade PDF that risk and compliance teams can attach to vendor risk assessments and AI Act conformity checks.

See what your agents are actually doing

Nexus gives you governance, observability and spend control across every agent you run.