← The journal
Technology

AI Usage and Cost Analytics: FinOps for Models and Agents

How to track AI spend by model, agent and workflow, set caps that act and route cost anomalies to an owner.

Swfte Journal / Technology

Cloud cost grew up with tagging, budgets and a FinOps discipline. Model spend arrived without most of it. It is metered in tokens, priced differently by model, and driven by behaviour that nobody in finance can see: a prompt that grew, an agent that loops, a workflow that calls a large model where a small one would do.

This post is about doing FinOps for models and agents: what you need to see, what you need to be able to control, and how a cost finding becomes a response with an owner. It also separates what the Swfte products measure and control today from what the Swfte Intelligence Platform is designed to add. We quote no savings figures, because we have none we can source here.

What makes AI cost different from cloud cost?

Four things.

It is behavioural. A virtual machine costs the same each hour. An agent's cost depends on how many steps it took, which model it chose and how much context it carried. Two runs of the same agent can differ widely.

It is spread across consumers. Spend is generated by agents, workflows, chat assistants and individual developers, often under one shared key. Without attribution, a bill is just a number.

Prices move. Models are repriced and replaced. A routing decision that was right in the spring may not be right now.

The unit that matters is the task. Cost per token is a supplier's number. What you care about is cost per completed piece of work, and whether the work was worth it.

What do you need to see?

Start with a breakdown along the dimensions that explain spend.

  • By model. Which models take the money, and is each justified for what it is used for?
  • By agent and by workflow. Which consumers drive it, and which step inside a workflow is the expensive one?
  • Over time. Is the trend steady, seasonal or a step change?
  • Against the previous period. What moved, and by how much?
  • Against a forecast. Where will you land if nothing changes?

The Swfte products measure this today. Usage and cost are available by model, by agent, by workflow and by workflow node, with time series and a list of top consumers. Per-request token and cost records support filtering, a breakdown, an export and an upstream health view. In Connect and Studio, dashboards cover spend by model, usage over time, a usage forecast, latency distribution, error breakdown, tool usage, a period comparison and a cost optimisation card. The Swfte command-line tool reports usage by model and costs. See Connect and cost optimisation and controls in Swfte Connect.

What do you need to control?

Seeing is half of it. The other half is limits that act before the bill does.

  • Caps. The products support workspace-level and per-model usage caps whose action on breach is to block, to downgrade to another model, or to alert.
  • Routing rules. Rules decide which model handles which traffic, so that a smaller, cheaper model takes work it can do and the large model is kept for work that needs it. We cover the idea in AI model routing and cost optimisation.
  • Savings view. A view of what rules and caps have saved, so you can check whether they are doing anything.

For wider discipline on controlling spend, read AI usage control and cost reduction.

What can you not do yet?

Be wary of any vendor that does not give you a list like this.

  • Attribution is by workspace, model, agent and workflow. There is no per-team view. A team is an organisational concept, and the products do not yet know who is on which team. The graph is designed to hold that.
  • Per-step cost on every worker trace is not available. It is on the list of things to join up.
  • You cannot yet ask in plain language why spend rose. A natural-language assistant exists for datasets, not for usage and cost telemetry.
  • Anomalies are computed on demand and not kept. Detection compares the current period with the previous one for usage, latency, errors and cost. There is no stored history of anomalies and no acknowledgement of them.

From a spend anomaly to a response: where the Intelligence Platform fits

Every FinOps programme hits the same wall. A dashboard shows a deviation. Someone has to work out whose it is, who can approve a change, and what a safe change looks like. That is a hunt through a directory, a spreadsheet and several inboxes, and it takes days.

The graph at the centre of the Intelligence Platform already resolves people, groups and reporting lines from Active Directory and LDAP, Microsoft Entra ID, Okta and Google Workspace, with accounts resolved into people. That is enough to find a group and an approver from a manager chain. Cost-centre and service ownership depend on collectors for cloud and business systems, which are on the roadmap.

The designed response is a FinOps agent, built from the finding. In the format we use for every governed agent:

  • Can: read usage and cost records for the teams it is scoped to; compare spend with the previous period; resolve the owning group and approver from the graph; draft a spend report and a proposed cap.
  • Cannot: change a budget or routing rule on its own; approve its own proposal; read unrelated personal data; pause another team's agents.
  • Requires approval: applying a cap or routing change; any message to a team outside its scope; any change to a budget.
  • Records: identity, usage data read, the owner it resolved with that fact's evidence status, model used, proposal, approver, action and outcome.

It starts at L1 Assist, then moves to L2 Approve once its proposals have been accepted on the record. The evidence status matters here. If the graph shows the owner as inferred or stale, the agent is built to say so and ask a person. It does not send a confident message to the wrong team.

How do you measure whether the response worked?

Because you started from a finding, you have a baseline. After a cap or a routing change, compare the same breakdown against the same period. Then write the result back, so that the next review starts from it. This is the measure and learn stages of the intelligence loop, described in the intelligence loop.

Include quality in the measure. A cap that cuts spend by degrading answers the business relied on is not a saving. Cost belongs beside the outcome, and the platform's views are designed to put them side by side.

A cost routine you can start this week

  1. Break it down. Spend by model, then by agent and workflow. Most surprises sit with a handful of consumers.
  2. Ask whether the model fits the task. Test a smaller model on your own evaluation, not on a general claim.
  3. Set the limit before the bill. A cap with a defined action on breach is cheaper than a retrospective.
  4. Name an owner for each large consumer. A cost nobody owns is a cost nobody reduces.
  5. Review a sample of runs. Look for loops and bloated context, which are the usual causes of runaway spend.
  6. Measure and write back. Compare against the baseline and record the result.

What is built, and what is not?

Built in the appliance: the temporal graph, directory sync with identity resolution, a local API, an outbound link with signed commands and a kill switch, and install packaging for a VM, Kubernetes and air-gapped environments. In progress: a pre-model sanitisation gateway. On the roadmap: additional collectors, a context-package API and MCP server, an action gateway with approvals, the production control plane, marketplace delivery, and the wiring into Cortex, Nexus, Studio and the Nexus harness. The unified usage and cost views, named dashboards, chart types, availability and pricing for the Intelligence Platform are not published yet.

Where to go next

Read the usage and cost analytics page, then AI observability. The wider models layer covers how models are chosen and governed. Swfte does not hold SOC 2 or ISO 27001 attestations and does not sign HIPAA business associate agreements today (a SOC 2 Type I audit is in preparation; see the trust page), and the platform is built for compliance-by-design, with the exact posture depending on your use case, jurisdiction, deployment and configuration. To talk through your own spend, talk to our team.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.