Platform / Intelligence / Usage and cost

AI usage and cost analytics: FinOps for models and agents

Know who is spending what on which model, set limits that act, and turn a spend anomaly into a governed response.

Model spend is the first AI cost that surprises a finance team, because it is metered by the request and spread across teams, agents and workflows that nobody has put in a single view. This page describes what the Swfte products already measure and control, what the Intelligence Platform is designed to add, and the limits we know about.

Why model spend needs its own discipline

Cloud infrastructure cost grew up with tagging, budgets and a FinOps practice. Model spend has arrived without most of that. It is metered in tokens, priced differently by model, and driven by behaviour that nobody on the finance side can see: a prompt that grew, an agent that loops, a workflow that calls a large model where a small one would do.

The result is a familiar pattern. Spend rises, someone notices a bill, and a manual hunt begins for the cause. By then the cause is several weeks old. A better position is one where usage is visible by model, agent and workflow as it happens, limits exist before the bill does, and a deviation reaches the person who can do something about it.

That last step is the one most tools stop short of, and the one the Intelligence Platform is designed to address.

What the Swfte products measure and control today

Described as built, from the products. Scope is a workspace, and the details are on the Connect and Studio product pages.

  • Usage and cost by model, agent and workflow

    Usage and cost broken down by model, and by agent, workflow and individual workflow node, with time series and a list of the top consumers.

  • Per-request records

    Per-request token and cost records with filtering, a breakdown view, a time series, an export and an upstream health view.

  • Dashboards

    In Connect and Studio, dashboards for spend by model, usage over time, a usage forecast, latency distribution, error breakdown, tool usage, a period-over-period comparison and a cost optimisation card.

  • Caps that act

    Workspace-level and per-model usage caps whose action on breach is to block, downgrade to another model, or alert.

  • Routing rules and savings

    Rules that decide which model handles which traffic, and a view of the savings those rules and caps produce.

  • Command-line access

    The Swfte command-line tool can report usage by model and costs, set a monthly spend cap, and list alerts and savings.

What it does not do yet

Honesty is part of cost analytics, so here are the gaps we know about. Attribution is by workspace, agent, model and workflow. We do not have a per-team view, and a team is an organisational concept that depends on knowing who is on it. That is exactly what the graph is for, and it is designed for rather than available.

Per-step cost on every worker trace is not yet joined to the trace itself. Natural-language questions about spend, such as why did costs rise in the support workflows last week, are not available. The assistant that turns questions into queries works over datasets and not over usage and cost telemetry.

Detected anomalies are computed on demand and not retained, and alert acknowledgement is available only in the administrative console. We list these here rather than let a design intent read as a feature.

What the Intelligence Platform is designed to add: owner, approver, action

A cost finding has three parts. What changed is the part every tool shows. Who owns it and who can approve a response are the parts that decide how fast anything happens, and those are questions about the organisation.

The graph already resolves people, groups and reporting lines from Active Directory and LDAP, Microsoft Entra ID, Okta and Google Workspace. That is enough to find a group and an approver from a manager chain. Cost-centre and service ownership depend on collectors for cloud and business systems, which are on the roadmap.

From there the design is a FinOps agent. It reads usage for the teams in scope, compares them with the previous period, resolves the owner and the approver, and drafts a report and a proposed cap. It proposes and does not apply. Applying a cap or a routing change requires approval, and every step is recorded: the data it read, the owner it resolved and that fact’s evidence status, the model it used, the proposal, the approver and the result. The worked example below sets it out in full.

A cost routine that closes the loop

Independent of any tool, and a good use of the loop.

  1. 01

    Break it down

    Look at spend by model, then by agent and workflow. Most surprises sit in one of a few consumers.

  2. 02

    Ask whether the model fits the task

    Routing rules and downgrades exist because a smaller model often does the job. Test it on your own evaluation, not on a general claim.

  3. 03

    Set the limit before the bill

    A cap with a clear action on breach is cheaper than a retrospective.

  4. 04

    Name an owner

    A cost nobody owns is a cost nobody reduces.

  5. 05

    Measure the response

    After a change, compare against the baseline and write the result back, so the next review starts with it.

Specifics we have not published

We state no savings figure and no benchmark, because we have none we can source here. The named Intelligence Platform dashboards are <named dashboards - founder to fill>. Availability is <availability - founder to fill> and pricing is <pricing - founder to fill>.

A worked example: spend anomaly to FinOps agent

Spend anomaly to FinOps agent

You see
Model spend for one team is running well above its own usual pattern.
The graph answers
Who owns this, and who can approve a cap?
It becomes
A FinOps agent that watches usage and cost, and proposes caps to the right person.

AI FinOps Agent

Can
  • Read usage and cost records for the teams it is scoped to
  • Compare spend with the previous period
  • Resolve the owning group and the approver from the graph
  • Draft a spend report and a proposed cap
Cannot
  • Change a budget or a routing rule on its own
  • Approve its own proposal
  • Read unrelated personal data
  • Pause another team’s agents
Requires approval
  • Applying a cap or a routing change
  • Any message to a team outside its scope
  • Any change to a budget
Records
  • Identity
  • Usage data read
  • Owner it resolved, with its evidence status
  • Model used
  • Proposal
  • Approver
  • Action
  • Outcome

Starts at: L1 Assist, then L2 Approve once its proposals have been accepted on the record.

Built today: People, groups and reporting lines, so the approver can be found from the manager chain.

Designed for: Service and cost-centre ownership, which depends on the collectors that are on the roadmap.

Usage and cost analytics

How usage and cost analytics connect to the closed loop

Cost is both a finding that starts a lap of the loop and a measure that ends it. A response that fixes a problem at too high a price is itself a finding.

The same eight stages are listed in order below.
  1. 01 · Layer 02Connect dataBring directory data today, and more systems over time, into a graph that lives in your environment.
  2. 02 · Layers 02 and 03AnalyseAsk questions in plain language, explore the graph, and look for trends and anomalies.(this page)
  3. 03 · Layers 02 and 03VisualiseSee the organisation, usage, agents and outcomes as maps, timelines, dashboards and evidence views.(this page)
  4. 04 · PeopleDecideChoose the response with the owner, the approver and the evidence status in front of you.
  5. 05 · Layers 04 to 06BuildTurn the insight into an agent, a workflow or a packaged solution.(this page)
  6. 06 · Trust FabricGovernIdentity, permissions, policy, audit and human approval apply while the thing runs.
  7. 07 · Layer 06MeasureTrack the outcome and the cost against the reason you built it.(this page)
  8. 08 · Layer 02LearnFeed what happened back into the graph, so the next question starts from more evidence.

Frequently asked questions

Can I see AI spend by agent and by model?

Yes. The products break usage and cost down by model, agent, workflow and workflow node, with time series and top consumers.

Can I stop spend automatically?

You can set workspace and per-model caps whose action on breach is to block, downgrade or alert. Applying a change proposed by an agent is designed to require human approval.

Is there a per-team view?

Not yet. Attribution is by workspace, model, agent and workflow. A per-team view depends on organisational data that the graph is designed to hold.

Can I ask why spend went up in plain language?

Not over usage and cost telemetry today. It is designed for, and the query language is not published yet.

Do you claim savings?

No. We state no savings figure. Measure the effect of a cap or a routing change against your own baseline.

Does cost data leave my environment?

Where the platform runs in your environment, customer data stays there by default. In connected mode only health counts, coverage summaries, diagnostics and command results go to Swfte.

Take usage and cost analytics further with Swfte

Start with one entry point. Add intelligence, agents, workflows and infrastructure as you prove value.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.