SecOps / Evidence

AI audit trail: evidence-grade records for agents and models

Record not just that an AI system acted, but who it was, what it read, which policy applied, who approved it and what happened next.

An audit log that says an API was called is not an audit trail for AI. When an agent acts, an auditor, an incident lead or a regulator wants the chain: the data it accessed, the model it used, the tools it called, the policy it was under, the approval it received and the outcome. That chain is traceability, and it is the evidence side of governance.

The problem: logs record events, not accountability

Application logs were designed for debugging. They capture that something happened, usually without the context that makes it explainable. For an AI agent the missing context is the point: the prompt that led to the action, the documents it retrieved, the model and its version, the policy that allowed it and the person who approved it. Without those, the question "why did the agent do that?" has no answer on record.

The gap hurts in three places. Security teams need it to investigate: a log without the causal chain turns a ten-minute question into a day of reconstruction. Auditors need it to test controls: they want to see that the approval rule fired, not that it exists on paper. And regulators increasingly expect it: the EU AI Act requires high-risk systems to technically allow automatic recording of events over their lifetime and expects deployers to keep the resulting logs for a period appropriate to the purpose, of at least six months unless other law says otherwise.

There is also a tension to manage. A trail detailed enough to prove what happened can contain personal or confidential data. Evidence-grade does not mean record everything; it means record the right things, protect them, keep them for a defined period and be able to show they were not altered.

What the trail records and how it is protected

Traceability on the platform is the chain data to model to agent to decision to action to outcome. Each link has a record, and the links are joined.

  • Identity and ownership

    Which agent or non-human identity acted, which version, which human or team owns it and which Trust Profile applied.

  • Inputs, with privacy tiers

    What the agent was asked and what it read. At the default privacy tier prompts are captured as fingerprint and length only, with no prompt text stored; deeper capture is a deliberate choice per data class.

  • Model and tools

    Which model and version produced the output, and every tool called, with target, arguments and result summary. Blocked calls are recorded with the policy reason.

  • Policy and approval

    Which policy set was evaluated, the verdict (allow, deny, warn, filter, escalate or require approval) and, where a person approved, who and when.

  • Action and outcome

    What changed: files, records, messages, installs. And what resulted. Every change is traceable back through the tool action to the prompt that caused it.

  • Protection and retention

    Records are written to a durable ledger and are designed to be append-only, with access control on who can read them. Retention follows your policy. <retention options and tamper-evidence mechanism - founder to fill>

Where it sits on the platform

One platform with several entry points. Each product below plays a defined part in this capability.

Platform products and the role each plays
Platform entry pointRole in this capability
NexusCaptures sessions, prompts (as fingerprints by default), tool actions, file changes and dependency installs to a durable ledger, and traces prompt to change.
Trust FabricAuditability, traceability and evidence as governance facets across all six layers.
BuildXLogs model calls through the gateway: which model, which request, what cost.
ConnectRecords calls to external systems and tools. Export to your SIEM is the route into existing detection and retention. <supported SIEM exports - founder to fill>
StudioRecords workflow runs, approvals and versions of the agent that ran.

What an audit agent can and cannot do

Auditing can be assisted by an agent that answers questions from the record. Its Trust Profile protects the integrity of that record.

Audit Query Agent

Answers auditors' and investigators' questions from the trail, with the records cited.

Audit Query Agent: what it can do, cannot do, requires approval for, and records
CanCannotRequires approvalRecords
  • Search and read the audit ledger within the requester's access
  • Rebuild the chain for a given action or incident
  • Summarise policy decisions and approvals for a period
  • Export a bounded evidence pack for a named reviewer
  • Write to, edit or delete audit records
  • Read prompt text where the privacy tier forbids it
  • Return records outside the requester's access
  • Alter its own logging
  • Exporting an evidence pack outside the organisation
  • Reading records at a deeper privacy tier
  • Any request spanning multiple business units
  • Requester and purpose
  • Records read and returned
  • Export created and recipient
  • Policy applied and approver

Controlled autonomy for evidence work

Reading the record is low-risk and can be highly automated. Changing it is not allowed at any level.

  1. L1 Assist

    The agent answers questions and shows the records behind each answer.

  2. L2 Approve

    The agent assembles evidence packs; a reviewer approves before they leave the system.

  3. L3 Supervise

    Scheduled evidence collection and exception reports run within limits and are monitored.

  4. L4 Autonomous

    Continuous control checks run independently and alert on gaps, within strict access bounds.

  5. L5 Adaptive

    Report formats and checks are refined from reviewer feedback under gated change. The record itself is never agent-writable.

The five-level model is described in full on the controlled autonomy page.

Frameworks the design is mapped to

Mappings of intent between controls and published risk lists. Not an assessment result.

Framework entries and how the controls are designed to address them
FrameworkEntryHow the design addresses it
OWASP Top 10 for LLM Applications (2025)LLM02 Sensitive Information DisclosurePrivacy tiers and access control on the trail stop the record becoming a second leak.
OWASP Top 10 for Agentic Applications (2026)ASI09 Human-Agent Trust Exploitation; ASI10 Rogue AgentsApprovals are recorded with the evidence shown to the approver, and activity from unregistered agents is visible.
MITRE ATLASTechnique labels on recorded behaviourRecorded events can be labelled with ATLAS techniques for detection and reporting. Confirm current IDs at atlas.mitre.org.

EU angle: record-keeping, oversight and accountability

The trail is where several duties meet. Swfte supports the evidence; your legal team decides what applies.

EU instruments and what Swfte supports evidence for
InstrumentReferenceSupports evidence for
EU AI ActArt. 12 record-keeping; Art. 14 human oversight; Art. 26 deployer log retentionSupports automatic event records over the system's lifetime, evidence of oversight actions and retained logs for deployers of high-risk systems, which must be kept at least six months unless other law provides otherwise.
GDPRArt. 5(2) accountability; Art. 32 security of processingSupports demonstrating the measures applied. Because logs can themselves contain personal data, retention and minimisation settings matter as much as capture.
NIS2Art. 21 risk-management measures; Art. 23 incident reportingSupports incident handling and the record behind each reporting stage.
DORAICT risk management and incident reportingSupports a classified, time-stamped record for financial entities in scope.

Swfte is built compliance-by-design. It provides the technical controls, governance mechanisms and evidence required to deploy AI within an organisation's applicable regulatory, security and policy requirements. The exact posture depends on the customer's use case, jurisdiction, deployment and configuration. Nothing on this page is legal advice. Swfte's own security attestations are listed on the trust page.

What this page does not claim

  • "Evidence-grade" describes what the record is designed to contain. It is not a statement that any court or regulator has accepted a Swfte record.
  • Tamper-evidence and retention mechanisms are described as design intent; the specifics are pending confirmation.
  • The trail covers what flows through the governed path. Activity outside it, such as an unmanaged tool, is not recorded, which is why discovery matters.
  • Nothing here claims that your use of the platform is compliant with any regulation. Swfte's own security attestations are listed on the trust page.

Frequently asked questions

What should an AI audit trail contain?

Identity and owner of the agent, what it was asked and read, the model and version, every tool call with target and result, the policy evaluated and its verdict, any human approval, the action taken and the outcome. The links need to be joined so the chain can be rebuilt.

Do you store our prompts?

At the default privacy tier Nexus captures prompts as a fingerprint and length only; no prompt text leaves the machine or is stored. Capturing more is an explicit choice, made per data class and governed by access control and retention.

How long are records kept?

As long as your policy says. Where the EU AI Act applies as a high-risk deployer, logs under your control must be kept for a period appropriate to the purpose and at least six months, unless other law, such as data protection law, provides otherwise. <retention options - founder to fill>

Can the audit trail be edited by an agent?

No. The record is designed to be append-only and is not writable by the agents it describes.

How does it connect to our SIEM?

Records are designed to be exportable so they feed existing detection and retention. <supported SIEM exports - founder to fill>

The rest of SecOps

Nine pages cover both halves. Each has a different job.

Build AI audit trail with Swfte

Start with one agent and one policy, or talk to the team about your environment, your data and your regulators.

See what your agents are actually doing

Nexus gives you governance, observability and spend control across every agent you run.