Every major LLM, scored against
NIST AI RMF and the EU AI Act.
Compliance scorecards for procurement and risk teams. Each model is audited against the four NIST AI RMF functions and EU AI Act Articles 10-15, with cited evidence from published documentation and live behavioural probes.
Explore the review framework01Purpose & scope
What decision will this assessment inform?
Name the model, intended use and the people affected. Keep a model evaluation separate from approval of a specific deployment.
02Evidence & method
Can someone else follow the assessment?
Record the methodology, source date and supporting material. Mark missing evidence explicitly.
03Ownership & review
Who owns the next decision?
Identify the reviewer, the unresolved questions and the conditions that would trigger reassessment.
The source library.
Snapshot · 2026-09-19. Read each scorecard for its cited evidence, dates and limitations.
Claude Opus 4.6 governance audit
Anthropic · Based on published documentation (0%) · Updated 2026-09-19
Claude Sonnet 4.6 governance audit
Anthropic · Based on published documentation (0%) · Updated 2026-09-19
Claude Haiku 4.5 governance audit
Anthropic · Based on published documentation (0%) · Updated 2026-09-19
GPT-5 governance audit
OpenAI · Based on published documentation (0%) · Updated 2026-09-19
GPT-4.5 governance audit
OpenAI · Based on published documentation (0%) · Updated 2026-09-19
o3-mini governance audit
OpenAI · Based on published documentation (0%) · Updated 2026-09-19
Gemini 2.5 Pro governance audit
Google · Based on published documentation (0%) · Updated 2026-09-19
Gemini 2.5 Flash governance audit
Google · Based on published documentation (0%) · Updated 2026-09-19
Llama 4 405B governance audit
Meta · Based on published documentation (0%) · Updated 2026-09-19
Llama 4 70B governance audit
Meta · Based on published documentation (0%) · Updated 2026-09-19
Mistral Large 2 governance audit
Mistral · Based on published documentation (0%) · Updated 2026-09-19
Mistral Small 3 governance audit
Mistral · Based on published documentation (0%) · Updated 2026-09-19
DeepSeek V3 governance audit
DeepSeek · Based on published documentation (0%) · Updated 2026-09-19
DeepSeek R1 governance audit
DeepSeek · Based on published documentation (0%) · Updated 2026-09-19
Qwen 3 governance audit
Alibaba · Based on published documentation (0%) · Updated 2026-09-19
Command R+ governance audit
Cohere · Based on published documentation (0%) · Updated 2026-09-19
Kimi K2 governance audit
Moonshot · Based on published documentation (0%) · Updated 2026-09-19
Grok 3 governance audit
xAI · Based on published documentation (0%) · Updated 2026-09-19
Jamba 1.5 governance audit
AI21 · Based on published documentation (0%) · Updated 2026-09-19
Phi-4 governance audit
Microsoft · Based on published documentation (0%) · Updated 2026-09-19
Gemma 3 governance audit
Google · Based on published documentation (0%) · Updated 2026-09-19
No references match. Try another title or topic.
21 references
What the audit answers.
What is an LLM governance audit?
A governance audit scores a large language model against a formal framework (NIST AI RMF, the EU AI Act, ISO/IEC 42001, or sector-specific rules) and produces a procurement-grade scorecard. The output covers the four NIST functions (Govern, Map, Measure, Manage) and EU AI Act Articles 10-15: data governance, technical documentation, transparency, human oversight, accuracy, robustness, and cybersecurity.
Which models are scored?
Every major frontier LLM available via API or open weights: Claude Opus 4.7 and Sonnet 4 (Anthropic), GPT-5.5 Pro and GPT-5.5 (OpenAI), Gemini 3.1 Pro and 3.0 (Google), DeepSeek V4 Pro and V4 (DeepSeek), Llama 4 (Meta), Grok 4 (xAI), Qwen 3 (Alibaba), Mistral Large, Kimi K2.5, and Command R+. New releases are added within 30 days of GA.
How are scores calculated?
Each model is scored on cited evidence from published documentation (model cards, system cards, transparency reports) plus live behavioural probes against the audit harness. Completeness reflects how much of the NIST + EU AI Act control set has verifiable evidence; the score is a weighted aggregate across the controls.
How often are the audits refreshed?
Audits refresh on every model version bump and quarterly otherwise. Each scorecard shows updatedAt: the timestamp of the last evidence pass.
Can I use these scorecards in a procurement RFP?
Yes. That is the primary use case. Each scorecard exports as a procurement-grade PDF that risk and compliance teams can attach to vendor risk assessments and AI Act conformity checks.