Validate · Intermediate

How to audit AI systems

  • Time: Two to four weeks for one system, depending on how quickly evidence arrives
  • Cost: Staff time. No licences are needed; the tools below are spreadsheets, a ticket tracker and a short script.
  • Level: Intermediate
On this page
  1. Short answer
  2. Before you start
  3. Internal audit, external audit and certification are different jobs
  4. 1. Set the scope, objective and independence
  5. 2. Confirm the inventory and map the data flows
  6. 3. Choose the criteria you will judge against
  7. 4. Send the evidence request
  8. 5. Test governance, approvals and change control
  9. 6. Test the logs and the audit trail
  10. 7. Re-perform performance and security tests
  11. 8. Test human oversight in practice
  12. 9. Write findings that can be fixed
  13. 10. Follow up and set the next audit cycle
  14. A simple rating scale for findings
  15. Troubleshooting
  16. Verify it worked
  17. Next steps
  18. FAQ
  19. How Swfte can help
  20. Sources and last verified

Short answer

To audit an AI system, fix its boundary and inventory first, pick the criteria you will judge it against (your policy, ISO/IEC 42001, NIST AI RMF, EU AI Act articles), send an evidence request, then test the controls by sampling real records: logs, approvals, changes, test results and oversight. Report each finding with criteria, evidence, risk and owner, and re-test after fixes. An audit is evidence-based, not an opinion.

The steps at a glance

  1. Set the scope, objective and independence
  2. Confirm the inventory and map the data flows
  3. Choose the criteria you will judge against
  4. Send the evidence request
  5. Test governance, approvals and change control
  6. Test the logs and the audit trail
  7. Re-perform performance and security tests
  8. Test human oversight in practice
  9. Write findings that can be fixed
  10. Follow up and set the next audit cycle

Before you start

Who this is for

  • Internal audit, risk and compliance teams asked to look at AI for the first time.
  • Security and engineering leads preparing a system for an external assessment.
  • Consultants running a client engagement who need a repeatable structure.

Probably not for you if

  • Teams who need a certificate: certification of an AI management system is done by an accredited body. ISO/IEC 42006:2025 sets requirements for those bodies; this guide prepares you for them but does not replace them.
  • Anyone who only wants a technical model test: use how to validate your AI for that part.

Prerequisites

  • A named executive sponsor and written permission to read records, logs and configurations.
  • An inventory of AI systems in scope, or time to build one (see the AI estate inventory post linked in step 2).
  • A person independent of the team that built the system, or a clear written note of any conflict of interest.
  • Access to the system owner and to someone who can pull logs and change records.
Time
Two to four weeks for one system, depending on how quickly evidence arrives
Cost
Staff time. No licences are needed; the tools below are spreadsheets, a ticket tracker and a short script.
Skill
Audit or risk experience helps; the technical tests need an engineer on call

Estimates are ours, not measurements, and move with your hardware, data and network.

Internal audit, external audit and certification are different jobs

People use “AI audit” for three things. An internal audit is your own independent function checking your systems against your policy and chosen criteria. An external or client audit is a third party doing the same under an engagement letter. Certification is a formal conformity assessment: ISO/IEC 42006:2025 specifies additional requirements to ISO/IEC 17021-1 for bodies that audit and certify an artificial intelligence management system against ISO/IEC 42001.

The method below works for the first two and produces the kind of evidence a certification body will ask for. Say at the start which job you are doing. It changes who you report to and how independent you must be.

  1. Step 1Set the scope, objective and independence

    You end up with: A signed one-page engagement note: which system, which period, what question, who you report to.

    Pick one system, not “our AI”. Name its boundary: the model or models, the prompts, the retrieval data, the tools it can call, the people and systems it touches, and the environments (development, staging, production). A boundary that leaves out the tools or the data is the usual way an AI audit misses the real risk.

    Write the audit question in one line: “Does the claims-triage assistant operate under effective controls for data access, accuracy, change and human oversight between 1 January and 30 September?” Then set the period you will sample from and the date evidence closes.

    Record independence. If you helped build the system, say so and have someone else lead the testing. Agree who receives the report and what happens to critical findings. Get the access permissions in writing now; waiting for them in week three is the most common delay.

  2. Step 2Confirm the inventory and map the data flows

    You end up with: A list of the components and data flows in scope that the system owner has confirmed as complete.

    You cannot test what nobody has listed. Ask for the AI system inventory and check it against reality: cloud bills for model API spend, gateway logs, identity-provider app grants and a short survey of teams. NIST AI RMF GOVERN 1.6 describes an AI system inventory as an organised database of artefacts relating to an AI system or model, which may include system documentation, incident response plans and data dictionaries. Use that as a minimum for what each entry holds.

    For the system in scope, draw the data flow: where inputs come from, which model processes them, where outputs go, what is stored and for how long, and which third parties see the data. Note every third-party model, dataset and package; NIST MAP 4 and GOVERN 6 both ask for risks from third-party software and data to be covered.

    Ask the owner to confirm the map in writing. Differences between the owner’s map and what you find in logs and configurations are your first findings. The AI estate inventory post gives a sample schema, and how to detect shadow AI covers finding systems nobody told you about.

    Checked against: NIST AI RMF Playbook: GOVERN, NIST AI RMF Playbook: MAP

  3. Step 3Choose the criteria you will judge against

    You end up with: A criteria matrix: each control objective mapped to the source it comes from.

    An audit compares what exists with a criterion. Start with the organisation’s own AI policy, because that is what management has committed to. Then add external criteria that apply. Do not try to test everything: choose the ones that match the system’s risk and the reason for the audit.

    Keep the matrix to one row per control objective, with the source clause beside it, so every finding can quote the criterion it breaks.

    Common criteria for an AI audit
    SourceUse it forWhere to look
    Your AI policy and proceduresThe baseline: did you do what you saidPolicy, standards, approval records
    ISO/IEC 42001 (AI management system)A management-system audit, and preparation for certificationThe management system requirements and Annex A controls in the standard itself, which you must buy
    NIST AI RMF 1.0A risk-practice review organised by Govern, Map, Measure and ManagePlaybook categories such as GOVERN 1, MAP 2, MEASURE 2, MANAGE 1
    EU AI ActLegal duties for high-risk systems: Article 9 risk management, 12 and 19 logs, 14 oversight, 15 accuracy and cybersecurity, 26 deployer duties, 72 post-market monitoring, Annex IV documentationThe Act text; confirm first whether the system is high-risk
    OWASP Top 10 for LLM Applications (2025)Security testing scope: LLM01 prompt injection to LLM10 unbounded consumptiongenai.owasp.org

    Checked against: NIST: AI Risk Management Framework, IEC Webstore: ISO/IEC 42006:2025, OWASP Top 10 for LLM Applications 2025, European Commission AI Act Service Desk: timeline

  4. Step 4Send the evidence request

    You end up with: A numbered request list sent to the system owner with a due date for every item.

    Ask for evidence in one list so nothing is requested twice. Number each item, say what period it covers and who should provide it. Ask for records the system produces by itself (logs, change tickets, approvals) in preference to documents written for the audit.

    The list below is a starting point. EU AI Act Annex IV is a useful checklist even outside the law: it names the system description, the development process, monitoring and control, appropriateness of performance metrics, the risk management system, lifecycle changes, standards applied and the post-market monitoring plan.

    Evidence request list (starting point)
    AreaEvidence to request
    GovernanceAI policy; named system owner; risk assessment and tier; approval for go-live; committee minutes
    DataData sources and lawful basis; data dictionary; retention rules; access lists for training and retrieval data
    Model and promptsModel name and version in production; prompt files under version control; list of tools and their permissions
    TestingTest charter, eval set and results for the last release; safety and red-team results; judge calibration if used
    Change controlChange tickets with approvals for model, prompt, data and tool changes in the period
    LoggingLog configuration; retention setting; sample of raw logs; process for access to logs
    Human oversightWho oversees; their training; approval records for risky actions; stop procedure
    Incidents and monitoringIncident register; monitoring dashboards; user complaint log; post-market monitoring plan
    Third partiesContracts and data processing terms for model and data providers; sub-processor list

    Checked against: EU AI Act Annex IV

  5. Step 5Test governance, approvals and change control

    You end up with: A worksheet showing, for each sampled change, whether it was approved, tested and recorded as the procedure says.

    Start with accountability. NIST GOVERN 2 expects accountability structures in which the right teams and individuals are empowered, responsible and trained. Check that the named owner exists, knows they are the owner, and has authority to stop the system. Check that approval for go-live was given by someone with that authority before the date the system went live, not after.

    Then test change control by sampling. Take the list of every change in the period to the model version, prompts, retrieval data and tool permissions. Pick a sample across change types and across months, not only the easy ones. For each, check the ticket, the approval, the test result before release and whether the production version now matches what was approved.

    The test that finds the most is comparing the live configuration with the approved one. Ask an engineer to show the running model identifier, prompt version and tool list on screen, in front of you. Differences between that and the last approved change are findings.

    Checked against: NIST AI RMF Playbook: GOVERN

  6. Step 6Test the logs and the audit trail

    You end up with: A conclusion on whether the system records enough to reconstruct what it did, and whether records are kept long enough.

    A system you cannot reconstruct cannot be audited. For each sampled interaction, ask the owner to show the trail from input to outcome: who or what made the request, which model and prompt version ran, what data was retrieved, which tools were called with which arguments, what was returned, and whether a human approved or changed anything.

    EU AI Act Article 12 says high-risk systems must technically allow automatic recording of events over the lifetime of the system, covering events relevant to risk situations, post-market monitoring and operation. Article 19 requires providers to keep the logs they control for a period appropriate to the intended purpose, at least six months unless other law says otherwise, and Article 26(6) sets the same minimum for deployers. Check the retention setting against that, and against your own policy and data protection duties; keeping personal data in logs longer than needed is its own finding.

    Use a script to draw a reproducible sample from an exported log, then trace each selected record end to end. The script below fixes the random seed so the sample can be repeated. Record any interaction where a link in the chain is missing.

    sample_events.py: reproducible random sample from a JSON Lines export · python
    import json
    import random
    
    SEED = 20261006   # write the seed in the workpaper so the sample can be repeated
    SAMPLE_SIZE = 25  # set by your audit methodology, not by this script
    
    with open("events.jsonl", encoding="utf-8") as f:
        events = [json.loads(line) for line in f if line.strip()]
    
    random.seed(SEED)
    sample = random.sample(events, k=min(SAMPLE_SIZE, len(events)))
    
    print(f"population={len(events)} sample={len(sample)} seed={SEED}")
    for e in sample:
        print(e.get("id", "<no id>"), e.get("timestamp", "<no timestamp>"))
    Run it · bash
    python3 sample_events.py

    Output shape

    population=<events in export> sample=<sample size> seed=20261006
    <event id> <timestamp>
    ...

    Checked against: EU AI Act Article 12, EU AI Act Article 19, EU AI Act Article 26

  7. Step 7Re-perform performance and security tests

    You end up with: Your own results for a small set of tests, compared with the results the team reported.

    Do not accept test results on trust. Ask for the eval set and the configuration, then re-run a subset yourself and compare. If the system owner’s reported pass rate cannot be reproduced within a reasonable range, that is a finding about the testing process, whatever the score. The method for building and judging those tests is in how to validate your AI.

    For security, scope the testing from the OWASP Top 10 for LLM Applications: prompt injection, sensitive information disclosure, supply chain, data and model poisoning, improper output handling, excessive agency, system prompt leakage, vector and embedding weaknesses, misinformation and unbounded consumption. Review the red-team report if one exists, check that fixes were re-tested, and run a few probes of your own. See how to red team an LLM.

    EU AI Act Article 15 asks high-risk systems to be resilient against attempts by third parties to alter their use, outputs or performance by exploiting vulnerabilities, and it names data poisoning, adversarial examples and model flaws. Use that as a prompt for what to ask about.

    Checked against: OWASP Top 10 for LLM Applications 2025, EU AI Act Article 15

  8. Step 8Test human oversight in practice

    You end up with: A conclusion on whether people can really understand, override and stop the system, supported by records.

    Oversight on paper is easy. Test it by looking for the traces. Article 14(4) of the EU AI Act lists what a person assigned to oversight must be enabled to do: understand the system’s capacities and limitations, remain aware of automation bias, interpret its output, decide not to use it or to override or reverse its output, and intervene or stop it. Article 26(2) says deployers assign oversight to people with the necessary competence, training and authority.

    Ask for training records for the overseers. Then look at approval and override records: if reviewers approve every item in seconds, or never override anything, oversight may be a rubber stamp. Ask someone to demonstrate the stop procedure. For agents, check that actions listed as needing approval cannot run without it. See how to set up human approval for AI agents.

    Checked against: EU AI Act Article 14, EU AI Act Article 26

  9. Step 9Write findings that can be fixed

    You end up with: A report in which each finding states the criterion, the evidence, the effect, the rating and a proposed action.

    Write each finding in five parts: criterion (what should be true, with the clause), condition (what you found, with the evidence reference), cause (why, if you know), effect (what could go wrong) and recommendation. Use the rating scale agreed at the start. Quote evidence by reference number, not by memory.

    Share the draft with the system owner before the report is final and record their response next to each finding. Factual errors get corrected; disagreement about risk is recorded, not erased. Be specific about what you did not test and why. An audit that implies full coverage is worse than one that states its limits.

    Agree an owner and a date for each action, then re-test before closing. Closing a finding on a promise is how the same issue appears in the next report.

  10. Step 10Follow up and set the next audit cycle

    You end up with: A tracked action list and a date for the next review, tied to changes in the system.

    Track each action to closure and re-test the control, not only the document. Feed confirmed gaps into the inventory and the risk register. Set triggers for an earlier review: a new model, a new tool with write access, a new data source or a serious incident.

    Keep the working papers with the report: criteria matrix, evidence list, sample records, scripts and seeds, interview notes and the management responses. A reviewer should be able to follow your conclusions back to evidence a year later.

A simple rating scale for findings

Use a scale the sponsor agrees to before fieldwork. This one is common in practice; adapt it to your own risk framework.

Finding ratings
RatingMeaningTypical response
CriticalA control is missing and the system can cause serious harm or a legal breach nowStop or restrict the system; fix before further use
MajorA control exists on paper but does not work in the sample testedFix within a stated period; owner and date recorded
MinorA control works but is incompletely documented or inconsistently appliedFix in normal course; track
ObservationAn improvement idea with no control failureOwner decides

Troubleshooting

What you seeLikely causeFix
The system owner says there are no logs, or the logs do not show prompts and tool calls.Logging was never enabled for AI-specific events, or it is turned off to save cost or protect privacy.Record it as a finding against the logging criterion, ask for a compensating source such as gateway or API-provider logs, and agree what minimum record will be kept going forward.
Evidence arrives late or incomplete.The request list was too long, or no one owns producing it.Ask the sponsor to name an evidence coordinator, split the list into must-have and nice-to-have, and state in the report which conclusions are limited by missing evidence.
The inventory lists one model but the logs show several.Teams are calling other models directly, or a gateway fallback routes to a model nobody approved.Treat it as a major finding on inventory and change control, check the gateway routing rules, and require every model in use to have an owner and an approval.
You cannot tell whether the system is high-risk under the EU AI Act.Use case and role (provider or deployer) are not documented.Document the intended purpose and role first, compare with the Annex III categories, and get legal advice for borderline cases. Audit against the stricter set of controls in the meantime.
The team’s test results cannot be reproduced.Unpinned model versions, a changed prompt, or an eval set that has moved.Ask for the exact configuration used. If it cannot be produced, report a testing-process finding and ask for a controlled re-run.
The model provider will not share documentation or testing evidence.Commercial confidentiality or a missing contract term.Record the limit, rely on your own testing of the deployed system, and recommend contract changes so future audits can obtain what they need.

Verify it worked

Next steps

Related guides

  • How to Validate Your AI: Eval Sets, Gates, Evidence: A system-level method to validate an AI product: define the task and risk, build a held-out eval set, score it, gate releases, sample for human review, monitor and keep an evidence pack.
  • How to Prepare for the EU AI Act: 2026 Readiness Steps: A readiness workflow for the EU AI Act: inventory AI systems, rule out prohibited uses, set provider or deployer role, classify risk, plan obligations and evidence, and track the dates after the Digital Omnibus.
  • How to Govern AI Agents: Identity, Policy, Approvals: Govern agents at runtime: list every agent, give each an identity and an owner, write down what it may and may not do in a Trust Profile, enforce allow, deny and approve rules, choose an autonomy level, record every action and review on a schedule.
  • How to Detect Shadow AI in Your Organisation: Define what counts as unsanctioned AI, then find it through identity grants, network logs, expense data and a short staff survey, rank what you find by the data it touches, replace the risky tools with approved ones, and keep monitoring in a proportionate way.
  • How to Do a DPIA for AI: GDPR Article 35 Steps: Decide whether an AI system needs a DPIA, then describe the processing, assess necessity, identify risks to people, choose measures, record the sign-off and review it, with AI-specific risks and a worked example.

Frequently asked questions

What is an AI audit?

An AI audit is an independent, evidence-based check of how an AI system is governed, built, tested, run and monitored, measured against criteria such as your own policy, ISO/IEC 42001, the NIST AI RMF or EU AI Act articles. It ends in findings with owners and dates, not in a pass mark alone.

How do you audit an AI system step by step?

Set scope and independence, confirm the inventory and data flows, choose criteria, send an evidence request, test governance and change control, test logs, re-perform performance and security tests, test human oversight, then write findings and follow up. Sample real records at each stage rather than relying on documents.

What is the difference between an AI audit and ISO 42001 certification?

An audit can be internal or by a client and judges a system against whatever criteria you choose. Certification is a third-party conformity assessment of an AI management system against ISO/IEC 42001 by a certification body, which must meet ISO/IEC 42006 requirements. An internal audit is good preparation for certification.

How long must AI system logs be kept under the EU AI Act?

For high-risk systems, providers and deployers must keep the automatically generated logs they control for a period appropriate to the intended purpose, of at least six months unless other Union or national law says otherwise (Articles 19 and 26(6)). Data protection law may limit how long personal data in logs can be kept.

Who can audit our AI systems?

Anyone with the right access, independence and competence: an internal audit function, a risk team separate from the builders, or an external firm. Competence needs both audit method and enough technical understanding to test logs, changes and evaluation results. Say in the report who did the work and any limits.

What evidence do auditors ask for?

The inventory and data flows, policy and risk assessment, approvals, version-controlled prompts and model identifiers, test charters and results, change tickets, log samples, oversight and training records, incident register and third-party contracts. Records produced by the system itself carry more weight than documents written for the audit.

How Swfte can help

You can run this audit with a spreadsheet and the evidence your own systems already produce. If you run AI through Swfte, these pages describe the records and controls the platform is designed to provide.

Swfte does not perform your audit and its platform does not make a system compliant by itself; the exact posture depends on your use case, jurisdiction, deployment and configuration. Which audit-evidence exports are available in your plan is <audit evidence export availability - founder to fill>.

Missing a step or found a command that no longer works? Tell us, or request a how-to.

Sources and last verified

Commands, versions and facts in this guide were checked against the sources below on . Tools change quickly: if something differs from what you see, trust the official documentation and let us know.

  1. NIST: AI Risk Management Framework: AI RMF 1.0 (NIST AI 100-1, 26 January 2023), voluntary use, four functions Govern, Map, Measure, Manage; Generative AI Profile NIST AI 600-1 (26 July 2024)
  2. NIST AI RMF Playbook: GOVERN: GOVERN 1 to 6 titles; GOVERN 1.6 inventory; GOVERN 2 accountability; GOVERN 6 third parties
  3. NIST AI RMF Playbook: MAP: MAP 1 to 4 titles including MAP 4 third-party software and data
  4. NIST AI RMF Playbook: MANAGE: MANAGE 1 to 4 titles
  5. IEC Webstore: ISO/IEC 42006:2025: Title, 7 July 2025 publication, scope: requirements for bodies auditing and certifying AI management systems per ISO/IEC 42001
  6. OWASP Top 10 for LLM Applications 2025: LLM01 to LLM10 titles
  7. EU AI Act Article 12: Automatic recording of events (logs) over the lifetime of high-risk systems
  8. EU AI Act Article 14: Human oversight and the abilities in Article 14(4)
  9. EU AI Act Article 15: Accuracy, resilience and cybersecurity; attacks named in 15(5)
  10. EU AI Act Article 19: Provider log retention of at least six months
  11. EU AI Act Article 26: Deployer duties: oversight assignment (26(2)), monitoring (26(5)), logs for at least six months (26(6))
  12. EU AI Act Annex IV: Technical documentation items 1 to 9
  13. European Commission AI Act Service Desk: timeline: Annex III high-risk rules apply 2 December 2027; Annex I 2 August 2028

Topics

  • audit
  • evidence
  • ISO 42001
  • NIST AI RMF
  • EU AI Act
  • logging

Machine-readable copies: this guide as markdown, index of all guides (JSON). Canonical address: https://www.swfte.com/how-to-audit-ai-systems.

See what your agents are actually doing

Nexus gives you governance, observability and spend control across every agent you run.