# How to audit AI systems

Canonical: https://www.swfte.com/how-to-audit-ai-systems
Last verified: 2026-10-06
Difficulty: Intermediate
Time: Two to four weeks for one system, depending on how quickly evidence arrives
Cost: Staff time. No licences are needed; the tools below are spreadsheets, a ticket tracker and a short script.

## Short answer

To audit an AI system, fix its boundary and inventory first, pick the criteria you will judge it against (your policy, ISO/IEC 42001, NIST AI RMF, EU AI Act articles), send an evidence request, then test the controls by sampling real records: logs, approvals, changes, test results and oversight. Report each finding with criteria, evidence, risk and owner, and re-test after fixes. An audit is evidence-based, not an opinion.

## Who this is for

- Internal audit, risk and compliance teams asked to look at AI for the first time.
- Security and engineering leads preparing a system for an external assessment.
- Consultants running a client engagement who need a repeatable structure.

Not for:
- Teams who need a certificate: certification of an AI management system is done by an accredited body. ISO/IEC 42006:2025 sets requirements for those bodies; this guide prepares you for them but does not replace them.
- Anyone who only wants a technical model test: use [how to validate your AI](https://www.swfte.com/how-to-validate-your-ai) for that part.

## Prerequisites

- A named executive sponsor and written permission to read records, logs and configurations.
- An inventory of AI systems in scope, or time to build one (see the AI estate inventory post linked in step 2).
- A person independent of the team that built the system, or a clear written note of any conflict of interest.
- Access to the system owner and to someone who can pull logs and change records.

## Internal audit, external audit and certification are different jobs

People use “AI audit” for three things. An internal audit is your own independent function checking your systems against your policy and chosen criteria. An external or client audit is a third party doing the same under an engagement letter. Certification is a formal conformity assessment: ISO/IEC 42006:2025 specifies additional requirements to ISO/IEC 17021-1 for bodies that audit and certify an artificial intelligence management system against ISO/IEC 42001.

The method below works for the first two and produces the kind of evidence a certification body will ask for. Say at the start which job you are doing. It changes who you report to and how independent you must be.

## Steps

### Step 1: Set the scope, objective and independence

Outcome: A signed one-page engagement note: which system, which period, what question, who you report to.

Pick one system, not “our AI”. Name its boundary: the model or models, the prompts, the retrieval data, the tools it can call, the people and systems it touches, and the environments (development, staging, production). A boundary that leaves out the tools or the data is the usual way an AI audit misses the real risk.

Write the audit question in one line: “Does the claims-triage assistant operate under effective controls for data access, accuracy, change and human oversight between 1 January and 30 September?” Then set the period you will sample from and the date evidence closes.

Record independence. If you helped build the system, say so and have someone else lead the testing. Agree who receives the report and what happens to critical findings. Get the access permissions in writing now; waiting for them in week three is the most common delay.

### Step 2: Confirm the inventory and map the data flows

Outcome: A list of the components and data flows in scope that the system owner has confirmed as complete.

You cannot test what nobody has listed. Ask for the AI system inventory and check it against reality: cloud bills for model API spend, gateway logs, identity-provider app grants and a short survey of teams. NIST AI RMF GOVERN 1.6 describes an AI system inventory as an organised database of artefacts relating to an AI system or model, which may include system documentation, incident response plans and data dictionaries. Use that as a minimum for what each entry holds.

For the system in scope, draw the data flow: where inputs come from, which model processes them, where outputs go, what is stored and for how long, and which third parties see the data. Note every third-party model, dataset and package; NIST MAP 4 and GOVERN 6 both ask for risks from third-party software and data to be covered.

Ask the owner to confirm the map in writing. Differences between the owner’s map and what you find in logs and configurations are your first findings. The [AI estate inventory post](https://www.swfte.com/blog/ai-estate-inventory-models-agents-data-flows) gives a sample schema, and [how to detect shadow AI](https://www.swfte.com/how-to-detect-shadow-ai) covers finding systems nobody told you about.

### Step 3: Choose the criteria you will judge against

Outcome: A criteria matrix: each control objective mapped to the source it comes from.

An audit compares what exists with a criterion. Start with the organisation’s own AI policy, because that is what management has committed to. Then add external criteria that apply. Do not try to test everything: choose the ones that match the system’s risk and the reason for the audit.

Keep the matrix to one row per control objective, with the source clause beside it, so every finding can quote the criterion it breaks.

**Common criteria for an AI audit**

| Source | Use it for | Where to look |
| --- | --- | --- |
| Your AI policy and procedures | The baseline: did you do what you said | Policy, standards, approval records |
| ISO/IEC 42001 (AI management system) | A management-system audit, and preparation for certification | The management system requirements and Annex A controls in the standard itself, which you must buy |
| NIST AI RMF 1.0 | A risk-practice review organised by Govern, Map, Measure and Manage | Playbook categories such as GOVERN 1, MAP 2, MEASURE 2, MANAGE 1 |
| EU AI Act | Legal duties for high-risk systems: Article 9 risk management, 12 and 19 logs, 14 oversight, 15 accuracy and cybersecurity, 26 deployer duties, 72 post-market monitoring, Annex IV documentation | The Act text; confirm first whether the system is high-risk |
| OWASP Top 10 for LLM Applications (2025) | Security testing scope: LLM01 prompt injection to LLM10 unbounded consumption | genai.owasp.org |

> NOTE: The EU AI Act timeline moved. The Commission service desk shows Annex III high-risk rules applying from 2 December 2027 and Annex I from 2 August 2028. Check [the timeline](https://www.swfte.com/eu/ai-act-timeline) before you assume a date, and take legal advice on whether a system is high-risk. The [EU AI Act high-risk checklist](https://www.swfte.com/eu/ai-act-high-risk-checklist) helps you scope it.

### Step 4: Send the evidence request

Outcome: A numbered request list sent to the system owner with a due date for every item.

Ask for evidence in one list so nothing is requested twice. Number each item, say what period it covers and who should provide it. Ask for records the system produces by itself (logs, change tickets, approvals) in preference to documents written for the audit.

The list below is a starting point. EU AI Act Annex IV is a useful checklist even outside the law: it names the system description, the development process, monitoring and control, appropriateness of performance metrics, the risk management system, lifecycle changes, standards applied and the post-market monitoring plan.

**Evidence request list (starting point)**

| Area | Evidence to request |
| --- | --- |
| Governance | AI policy; named system owner; risk assessment and tier; approval for go-live; committee minutes |
| Data | Data sources and lawful basis; data dictionary; retention rules; access lists for training and retrieval data |
| Model and prompts | Model name and version in production; prompt files under version control; list of tools and their permissions |
| Testing | Test charter, eval set and results for the last release; safety and red-team results; judge calibration if used |
| Change control | Change tickets with approvals for model, prompt, data and tool changes in the period |
| Logging | Log configuration; retention setting; sample of raw logs; process for access to logs |
| Human oversight | Who oversees; their training; approval records for risky actions; stop procedure |
| Incidents and monitoring | Incident register; monitoring dashboards; user complaint log; post-market monitoring plan |
| Third parties | Contracts and data processing terms for model and data providers; sub-processor list |

### Step 5: Test governance, approvals and change control

Outcome: A worksheet showing, for each sampled change, whether it was approved, tested and recorded as the procedure says.

Start with accountability. NIST GOVERN 2 expects accountability structures in which the right teams and individuals are empowered, responsible and trained. Check that the named owner exists, knows they are the owner, and has authority to stop the system. Check that approval for go-live was given by someone with that authority before the date the system went live, not after.

Then test change control by sampling. Take the list of every change in the period to the model version, prompts, retrieval data and tool permissions. Pick a sample across change types and across months, not only the easy ones. For each, check the ticket, the approval, the test result before release and whether the production version now matches what was approved.

The test that finds the most is comparing the live configuration with the approved one. Ask an engineer to show the running model identifier, prompt version and tool list on screen, in front of you. Differences between that and the last approved change are findings.

> TIP: Keep the sample selection reproducible. Record the method, the seed and the population size so another auditor gets the same sample.

### Step 6: Test the logs and the audit trail

Outcome: A conclusion on whether the system records enough to reconstruct what it did, and whether records are kept long enough.

A system you cannot reconstruct cannot be audited. For each sampled interaction, ask the owner to show the trail from input to outcome: who or what made the request, which model and prompt version ran, what data was retrieved, which tools were called with which arguments, what was returned, and whether a human approved or changed anything.

EU AI Act Article 12 says high-risk systems must technically allow automatic recording of events over the lifetime of the system, covering events relevant to risk situations, post-market monitoring and operation. Article 19 requires providers to keep the logs they control for a period appropriate to the intended purpose, at least six months unless other law says otherwise, and Article 26(6) sets the same minimum for deployers. Check the retention setting against that, and against your own policy and data protection duties; keeping personal data in logs longer than needed is its own finding.

Use a script to draw a reproducible sample from an exported log, then trace each selected record end to end. The script below fixes the random seed so the sample can be repeated. Record any interaction where a link in the chain is missing.

sample_events.py: reproducible random sample from a JSON Lines export:

```python
import json
import random

SEED = 20261006   # write the seed in the workpaper so the sample can be repeated
SAMPLE_SIZE = 25  # set by your audit methodology, not by this script

with open("events.jsonl", encoding="utf-8") as f:
    events = [json.loads(line) for line in f if line.strip()]

random.seed(SEED)
sample = random.sample(events, k=min(SAMPLE_SIZE, len(events)))

print(f"population={len(events)} sample={len(sample)} seed={SEED}")
for e in sample:
    print(e.get("id", "<no id>"), e.get("timestamp", "<no timestamp>"))
```

Run it:

```bash
python3 sample_events.py
```

Output shape:

```text
population=<events in export> sample=<sample size> seed=20261006
<event id> <timestamp>
...
```

### Step 7: Re-perform performance and security tests

Outcome: Your own results for a small set of tests, compared with the results the team reported.

Do not accept test results on trust. Ask for the eval set and the configuration, then re-run a subset yourself and compare. If the system owner’s reported pass rate cannot be reproduced within a reasonable range, that is a finding about the testing process, whatever the score. The method for building and judging those tests is in [how to validate your AI](https://www.swfte.com/how-to-validate-your-ai).

For security, scope the testing from the OWASP Top 10 for LLM Applications: prompt injection, sensitive information disclosure, supply chain, data and model poisoning, improper output handling, excessive agency, system prompt leakage, vector and embedding weaknesses, misinformation and unbounded consumption. Review the red-team report if one exists, check that fixes were re-tested, and run a few probes of your own. See [how to red team an LLM](https://www.swfte.com/how-to-red-team-an-llm).

EU AI Act Article 15 asks high-risk systems to be resilient against attempts by third parties to alter their use, outputs or performance by exploiting vulnerabilities, and it names data poisoning, adversarial examples and model flaws. Use that as a prompt for what to ask about.

### Step 8: Test human oversight in practice

Outcome: A conclusion on whether people can really understand, override and stop the system, supported by records.

Oversight on paper is easy. Test it by looking for the traces. Article 14(4) of the EU AI Act lists what a person assigned to oversight must be enabled to do: understand the system’s capacities and limitations, remain aware of automation bias, interpret its output, decide not to use it or to override or reverse its output, and intervene or stop it. Article 26(2) says deployers assign oversight to people with the necessary competence, training and authority.

Ask for training records for the overseers. Then look at approval and override records: if reviewers approve every item in seconds, or never override anything, oversight may be a rubber stamp. Ask someone to demonstrate the stop procedure. For agents, check that actions listed as needing approval cannot run without it. See [how to set up human approval for AI agents](https://www.swfte.com/how-to-set-up-human-approval-for-ai-agents).

### Step 9: Write findings that can be fixed

Outcome: A report in which each finding states the criterion, the evidence, the effect, the rating and a proposed action.

Write each finding in five parts: criterion (what should be true, with the clause), condition (what you found, with the evidence reference), cause (why, if you know), effect (what could go wrong) and recommendation. Use the rating scale agreed at the start. Quote evidence by reference number, not by memory.

Share the draft with the system owner before the report is final and record their response next to each finding. Factual errors get corrected; disagreement about risk is recorded, not erased. Be specific about what you did not test and why. An audit that implies full coverage is worse than one that states its limits.

Agree an owner and a date for each action, then re-test before closing. Closing a finding on a promise is how the same issue appears in the next report.

### Step 10: Follow up and set the next audit cycle

Outcome: A tracked action list and a date for the next review, tied to changes in the system.

Track each action to closure and re-test the control, not only the document. Feed confirmed gaps into the inventory and the risk register. Set triggers for an earlier review: a new model, a new tool with write access, a new data source or a serious incident.

Keep the working papers with the report: criteria matrix, evidence list, sample records, scripts and seeds, interview notes and the management responses. A reviewer should be able to follow your conclusions back to evidence a year later.

## A simple rating scale for findings

Use a scale the sponsor agrees to before fieldwork. This one is common in practice; adapt it to your own risk framework.

**Finding ratings**

| Rating | Meaning | Typical response |
| --- | --- | --- |
| Critical | A control is missing and the system can cause serious harm or a legal breach now | Stop or restrict the system; fix before further use |
| Major | A control exists on paper but does not work in the sample tested | Fix within a stated period; owner and date recorded |
| Minor | A control works but is incompletely documented or inconsistently applied | Fix in normal course; track |
| Observation | An improvement idea with no control failure | Owner decides |

## Troubleshooting

| Symptom | Likely cause | Fix |
| --- | --- | --- |
| The system owner says there are no logs, or the logs do not show prompts and tool calls. | Logging was never enabled for AI-specific events, or it is turned off to save cost or protect privacy. | Record it as a finding against the logging criterion, ask for a compensating source such as gateway or API-provider logs, and agree what minimum record will be kept going forward. |
| Evidence arrives late or incomplete. | The request list was too long, or no one owns producing it. | Ask the sponsor to name an evidence coordinator, split the list into must-have and nice-to-have, and state in the report which conclusions are limited by missing evidence. |
| The inventory lists one model but the logs show several. | Teams are calling other models directly, or a gateway fallback routes to a model nobody approved. | Treat it as a major finding on inventory and change control, check the gateway routing rules, and require every model in use to have an owner and an approval. |
| You cannot tell whether the system is high-risk under the EU AI Act. | Use case and role (provider or deployer) are not documented. | Document the intended purpose and role first, compare with the Annex III categories, and get legal advice for borderline cases. Audit against the stricter set of controls in the meantime. |
| The team’s test results cannot be reproduced. | Unpinned model versions, a changed prompt, or an eval set that has moved. | Ask for the exact configuration used. If it cannot be produced, report a testing-process finding and ask for a controlled re-run. |
| The model provider will not share documentation or testing evidence. | Commercial confidentiality or a missing contract term. | Record the limit, rely on your own testing of the deployed system, and recommend contract changes so future audits can obtain what they need. |

## Verify it worked

- [ ] The engagement note names the system, boundary, period, question, sponsor and any conflict of interest.
- [ ] Every control tested is traceable to a criterion in the matrix.
- [ ] Each sample can be repeated: method, population size and seed are recorded.
- [ ] You compared the live configuration with the last approved change on screen, not from a document.
- [ ] You traced at least ten interactions from input to outcome and recorded gaps.
- [ ] Every finding has criterion, condition, cause, effect, rating, owner and date, and the owner has responded.
- [ ] The report states what was not tested.
- [ ] A re-test date exists for each action and for the next audit trigger.

## Next steps

- [How to validate your AI](https://www.swfte.com/how-to-validate-your-ai): Build the test evidence an auditor will ask to see.
- [How to prepare for the EU AI Act](https://www.swfte.com/how-to-prepare-for-the-eu-ai-act): Turn audit gaps into a compliance plan.
- [How to govern AI agents](https://www.swfte.com/how-to-govern-ai-agents): Set the controls an audit of agents will test.
- [AI audit trail](https://www.swfte.com/secops/ai-audit-trail): See what a useful audit record contains.
- [The AI estate inventory post](https://www.swfte.com/blog/ai-estate-inventory-models-agents-data-flows): A sample schema for the inventory in step 2.

## FAQ

### What is an AI audit?

An AI audit is an independent, evidence-based check of how an AI system is governed, built, tested, run and monitored, measured against criteria such as your own policy, ISO/IEC 42001, the NIST AI RMF or EU AI Act articles. It ends in findings with owners and dates, not in a pass mark alone.

### How do you audit an AI system step by step?

Set scope and independence, confirm the inventory and data flows, choose criteria, send an evidence request, test governance and change control, test logs, re-perform performance and security tests, test human oversight, then write findings and follow up. Sample real records at each stage rather than relying on documents.

### What is the difference between an AI audit and ISO 42001 certification?

An audit can be internal or by a client and judges a system against whatever criteria you choose. Certification is a third-party conformity assessment of an AI management system against ISO/IEC 42001 by a certification body, which must meet ISO/IEC 42006 requirements. An internal audit is good preparation for certification.

### How long must AI system logs be kept under the EU AI Act?

For high-risk systems, providers and deployers must keep the automatically generated logs they control for a period appropriate to the intended purpose, of at least six months unless other Union or national law says otherwise (Articles 19 and 26(6)). Data protection law may limit how long personal data in logs can be kept.

### Who can audit our AI systems?

Anyone with the right access, independence and competence: an internal audit function, a risk team separate from the builders, or an external firm. Competence needs both audit method and enough technical understanding to test logs, changes and evaluation results. Say in the report who did the work and any limits.

### What evidence do auditors ask for?

The inventory and data flows, policy and risk assessment, approvals, version-controlled prompts and model identifiers, test charters and results, change tickets, log samples, oversight and training records, incident register and third-party contracts. Records produced by the system itself carry more weight than documents written for the audit.

## How Swfte can help

You can run this audit with a spreadsheet and the evidence your own systems already produce. If you run AI through Swfte, these pages describe the records and controls the platform is designed to provide.

- [AI audit trail](https://www.swfte.com/secops/ai-audit-trail): What an audit record for AI activity should hold.
- [Platform governance](https://www.swfte.com/platform/governance): Policies, approvals and evidence in one layer.
- [Trust Profile](https://www.swfte.com/platform/trust-profile): A per-system record of owner, risk, data and approvals.
- [AI governance](https://www.swfte.com/ai-governance): The wider governance picture.

Swfte does not perform your audit and its platform does not make a system compliant by itself; the exact posture depends on your use case, jurisdiction, deployment and configuration. Which audit-evidence exports are available in your plan is <audit evidence export availability - founder to fill>.

## Sources

- [NIST: AI Risk Management Framework](https://www.nist.gov/itl/ai-risk-management-framework): AI RMF 1.0 (NIST AI 100-1, 26 January 2023), voluntary use, four functions Govern, Map, Measure, Manage; Generative AI Profile NIST AI 600-1 (26 July 2024)
- [NIST AI RMF Playbook: GOVERN](https://airc.nist.gov/airmf-resources/playbook/govern/): GOVERN 1 to 6 titles; GOVERN 1.6 inventory; GOVERN 2 accountability; GOVERN 6 third parties
- [NIST AI RMF Playbook: MAP](https://airc.nist.gov/airmf-resources/playbook/map/): MAP 1 to 4 titles including MAP 4 third-party software and data
- [NIST AI RMF Playbook: MANAGE](https://airc.nist.gov/airmf-resources/playbook/manage/): MANAGE 1 to 4 titles
- [IEC Webstore: ISO/IEC 42006:2025](https://webstore.iec.ch/en/publication/108460): Title, 7 July 2025 publication, scope: requirements for bodies auditing and certifying AI management systems per ISO/IEC 42001
- [OWASP Top 10 for LLM Applications 2025](https://genai.owasp.org/llm-top-10/): LLM01 to LLM10 titles
- [EU AI Act Article 12](https://artificialintelligenceact.eu/article/12/): Automatic recording of events (logs) over the lifetime of high-risk systems
- [EU AI Act Article 14](https://artificialintelligenceact.eu/article/14/): Human oversight and the abilities in Article 14(4)
- [EU AI Act Article 15](https://artificialintelligenceact.eu/article/15/): Accuracy, resilience and cybersecurity; attacks named in 15(5)
- [EU AI Act Article 19](https://artificialintelligenceact.eu/article/19/): Provider log retention of at least six months
- [EU AI Act Article 26](https://artificialintelligenceact.eu/article/26/): Deployer duties: oversight assignment (26(2)), monitoring (26(5)), logs for at least six months (26(6))
- [EU AI Act Annex IV](https://artificialintelligenceact.eu/annex/4/): Technical documentation items 1 to 9
- [European Commission AI Act Service Desk: timeline](https://ai-act-service-desk.ec.europa.eu/en/ai-act/timeline/timeline-implementation-eu-ai-act): Annex III high-risk rules apply 2 December 2027; Annex I 2 August 2028

Last verified against these sources on 2026-10-06.
