AI Observability for Agents and Workflows: What to Capture
What AI observability must answer beyond latency: which model, data, tool, policy and approver, at what cost.
A service either answered or it did not. An AI agent answered, but the interesting question is how: on whose authority, from which data, using which model, calling which tools, at what cost, under which policy, with or without a person in the loop. If your monitoring cannot answer those questions, it is monitoring the plumbing and missing the system.
This post sets out what AI observability for agents and workflows has to capture, which parts of that the Swfte products capture today, and where the Swfte Intelligence Platform is designed to go further. We try to keep the line between built and designed-for clear, because observability is the area where overclaiming does the most harm.
What is different about observing an agent?
Three things.
The unit of work is a chain, not a request. An agent run involves a prompt, a plan, several model calls, tool calls, retrieved context and often a human decision. A single latency number hides all of it.
The cost is variable and behavioural. A web request costs roughly the same each time. An agent's cost depends on how many steps it took, which model it chose and how much context it carried. Two runs of the same agent can differ by an order of magnitude.
The risk is in the content. An agent that read a document containing an instruction, or that touched data outside its scope, has done something a status code will never show. The record has to contain enough to reconstruct what it read and what it did.
The platform's name for the whole chain is traceability: data, model, agent, decision, action, outcome. Observability is how you look at it day to day.
What should you capture for every run?
Ten fields, which match the record the platform keeps for governed agents.
- Identity. Which agent or workflow, run as which identity, on whose behalf.
- Data accessed. What it read, from which systems, and the classification of each.
- Model used. Which model and which version, and whether a router chose it.
- Output. What it produced, subject to your retention and privacy policy.
- Tools called. Each call, its target, and whether it was allowed.
- Policy applied. Which rule fired and with what result: allow, deny, warn, filter, escalate or require human approval.
- Decision. What the agent decided and on what basis.
- Approval. Who approved, when, and on what evidence.
- Action. What actually happened in the world.
- Outcome. What came of it, and whether it was the intended effect.
Cost sits alongside all ten: tokens, model price and the total for the run.
What do the Swfte products capture today?
The products capture different slices of that list, and we describe each as it is.
- Model request records. Connect and Studio keep per-request records of model usage, tokens and cost, with a detail view for a single request.
- Workflow execution traces. Studio shows traces for a workflow run, with a per-node panel so you can see which step did what.
- Worker traces. Autonomous workers produce append-only traces with an integrity check and a live stream while they run. Per-step cost is not yet joined to those traces, and we say so.
- Agent capture and policy. Nexus captures agent actions, enforces policy in-flight and traces agents and connections. It is deepest today for coding agents.
- Audit and approvals. Audit events can be exported, and Nexus has an approval queue where a person decides on a pending action.
What do the Swfte products not capture yet?
Gaps are information. Here are the ones we know about.
- Per-step cost on every worker trace is not available.
- Detected anomalies are computed on demand and not kept as history, and anomaly acknowledgement is not available.
- Natural-language questions over usage, cost and trace telemetry are not available. A natural-language assistant exists for datasets, and it does not cover telemetry.
- Nothing automatically turns a detected issue into a built agent or workflow. A worker template for triaging an issue exists, and nothing connects a detected issue to it.
- The issue and struggle detection that Swfte runs covers sessions in its own web applications. It does not trace AI agent runs, and we do not offer it as such.
How does observability become a loop?
An observation that ends in a dashboard is the same failure as an insight that ends in a chart. The point of looking is to change something. In the platform's loop, observability supplies evidence for three stages: analysing what happened, visualising it, and measuring whether a response worked. What turns it into a loop is the stages in between: decide, build and govern.
The Intelligence Platform is designed to put traces, spend and policy decisions on the same graph as people, groups and reporting lines, so that an observation can name an owner and an approver. Today the graph resolves people, groups and reporting lines from Active Directory and LDAP, Microsoft Entra ID, Okta and Google Workspace. Service and system ownership depend on collectors that are on the roadmap.
From there, the response is designed to start from the observation. Take policy denials rising in one part of the organisation. The graph names the agents and owners involved. A SecOps response agent, at L1 Assist, groups the decisions, pulls each agent's Trust Profile and recent actions, opens an incident with the evidence attached and notifies the owner. Containment and any policy change require approval. SecOps agents are in beta, so we describe that as design intent. See SecOps agents and the AI incident response runbook.
A routine that works whatever tools you have
Four habits.
Define a bad run before you need to. An unexpected tool, a policy denial, a cost outside a range, an approval skipped, a reply outside scope.
Read a sample, not only the alerts. Alerts catch what you thought of. A weekly read of a random sample of runs catches what you did not.
Keep the record tamper-evident. A log that someone can edit is not evidence. The Intelligence Platform's audit log is hash-chained, with signed checkpoints, so a break is detectable.
Close the loop. Every recurring finding should end in a change: a policy, an approval rule, a new agent or a retired one.
For prompt-level analytics see our guide to LLM observability and prompt analytics, and for monitoring patterns across many agents see AI agent clusters and behaviour monitoring.
Does observability itself need governing?
Yes, and this is often overlooked. Traces contain prompts, retrieved documents and outputs, which can include personal and sensitive data. Observability data needs classification, retention and access control like anything else, and an analyst looking at a trace should see only what they are allowed to see.
In the appliance, personal and restricted data are local-only by default, secrets and credentials are never stored, and reads are filtered by what the asking person may see. Retention can be set so that personal data of deleted accounts is replaced with a pseudonym after a period you choose.
What is built, and what is not?
Built in the appliance: the temporal graph with insert-only history and as-of reads, directory sync with identity resolution, a local API, the outbound link with signed commands and a kill switch, and install packaging. In progress: a pre-model sanitisation gateway. On the roadmap: additional collectors, a context-package API and MCP server, an action gateway with approvals, the production control plane, marketplace delivery and the wiring into Cortex, Nexus, Studio and the Nexus harness. The unified views are design intent. We give no dates.
Where to go next
Read the page on AI observability, then usage and cost analytics. For the governance background see AI governance. Swfte does not hold SOC 2 or ISO 27001 attestations and does not sign HIPAA business associate agreements today (a SOC 2 Type I audit is in preparation; see the trust page), and the platform is built for compliance-by-design; the exact posture depends on your use case, jurisdiction, deployment and configuration. To discuss your own observability, talk to our team.