Updated Oct 6, 2026 · 8 min read

AI FinOps (October 2026)

TL;DR: AI FinOps is financial accountability for token-metered AI spend: see it, attribute it, cut the waste and set limits. The price of the same work varies widely by model, so start with a price index, then add attribution and budgets.

Why AI spend behaves differently

The meter is the token, not the server

Most LLM APIs bill input and output tokens separately. Spend follows how much text goes in and comes out, so it moves with prompt length, retries and agent loop depth rather than with provisioned capacity.

Output costs more than input

Output tokens are priced higher than input tokens on almost every hosted model. Long answers, reasoning tokens and verbose tool output are where bills grow.

Model choice swings the price by a large factor

The same prompt can cost very different amounts depending on the model. The table below shows the spread in our price index today.

Allocation is harder than for cloud servers

The consumer of a model’s output is often several layers away from the billed key. Without attribution at call time, the invoice shows a total and nothing about who or what caused it.

The spread in our price index

The AI model price index lists input and output prices per million tokens for 63 hosted models (open-weight models you self-host have no token price and are left out here). Across those models the median output price is 4.1x the input price. The blended figure below is the simple average of the input and output price.

ModelInput / 1MOutput / 1MBlended / 1M
Lowest blended price
DeepSeek V4 Flash$0.14$0.28$0.21
Gemini 2.0 Flash$0.10$0.40$0.25
Llama 4 Scout$0.15$0.40$0.28
Highest blended price
GPT-5.5 Pro$30.00$180.00$105.00
Claude Opus 4$15.00$75.00$45.00
Claude Fable 5.1$10.00$50.00$30.00

Prices are provider list prices as recorded in the index and change often. Cheaper is not better by itself: check quality on your own prompts, and see the true-cost guide for what the headline price leaves out.

Inform, optimize, operate

The FinOps Foundation applies its existing framework to AI. These are its three phases, read for token-metered spend.

Inform

See token use and cost, and attribute it to the team, product and model responsible. Tag every call with at least the model, the use case and the owner.

Optimize

Cut waste: match model size to task, cache repeated prefixes, batch work that can wait, trim context, and route to a cheaper model where quality allows.

Operate

Set quotas, budgets and alerts, review spend on a cadence, and revise forecasts more often than you would for provisioned infrastructure.

Where Swfte fits

  • Connect routes model calls through one endpoint with cost analytics, budgets and alerts, and failover between providers.
  • Nexus records token usage and cost for coding agents and rolls it up by user, repo, terminal and model.
  • The price index lets you compare models before you route to them. Which controls apply depends on your configuration.

Frequently asked questions

What is AI FinOps?

AI FinOps is the practice of applying FinOps (financial accountability for variable technology spend) to AI usage: tracking token and inference cost, attributing it to teams and products, optimizing it, and setting policies that keep it under control. The FinOps Foundation applies its existing framework, with its Inform, Optimize and Operate phases, to AI rather than defining a separate one.

Why is LLM spend hard to forecast?

It scales with tokens, not provisioned capacity. Prompt length, retries, reasoning tokens and agent loops all change the bill, and new models change prices, so forecasts need to be revised more often than for traditional cloud spend.

What should I measure first?

Cost per token by model, and spend by team, product and use case. The FinOps Foundation defines a token cost KPI as total cost divided by tokens used. Add input and output tokens separately, because they are priced differently.

How does Swfte help with AI spend?

Connect routes model calls through one endpoint with cost analytics, budgets and alerts. Nexus records token usage and cost for coding agents and rolls it up by user, repo, terminal and model. Our price index lists input and output prices for every model so you can compare before you route. Which controls apply depends on your configuration.

Where do the prices on this page come from?

From the AI model price index on this site, which lists each model’s input and output price per million tokens from its provider’s published pricing. Prices change, so check the index and the provider before you budget.

Find out where your AI spend goes

The free AI usage risk audit looks at usage metadata under NDA, and a free Connect key lets you route one workload through the gateway and compare cost.

Metadata only · NDA first · Design-partner terms on request