← The journal
Working Theory

Why Frontier AI Got Cheaper in September 2026: Our Working Theory on the Cuts

Anthropic and OpenAI cut frontier prices in a week and never said why. Our theory, and what would disprove it.

Swfte Journal / Working Theory

A CTO opens the model bill on the morning of 30 September, and the CFO has one question. Why did it get cheaper, and will it stay that way?

The honest answer starts with two facts. Between 22 and 29 September, Anthropic priced Claude Opus 5.5 at $4 in, $20 out per million tokens, against $5 and $25 for Opus 5. OpenAI priced GPT-6 Sol at $2 and $10, which VentureBeat calculated as exactly half of GPT-5.6 Sol. Neither company told us how it got there.

This post is our working theory: recent research plus low-level kernel and serving work explains a large part of why frontier AI is getting cheaper. It is a theory. We keep a hard line between what a vendor confirmed and what we infer, and we finish with what would prove us wrong. The launches themselves are covered in Claude Opus 5.5 and Sonnet 5.5 and GPT-6 Sol, Luna and 6.1 Sol.

Confirmed: cheaper to serve, fewer tokens, no mechanism

Confirmed means a vendor said it on the record. Our inference means we are reasoning from public research. A sentence with neither label is arithmetic.

What Anthropic confirmed on its Opus 5.5 page: "Opus 5.5 requires less compute to serve than Opus 5, and its pricing reflects that." And: "It costs less per token than Opus 5 and uses fewer tokens per task, which nets out to a 40% drop in costs." For Sonnet 5.5 the per-token price did not move, still $2 and $10; Anthropic credits fewer tokens per task and batched tool calls.

Anthropic also said cache reads "make up the majority of agentic and coding work costs", and priced them at $0.20 for Opus 5.5, 60% below Opus 5. Fable 5.1 sits at 0.025 times the input rate. No explanation is attached.

What OpenAI confirmed, as reported by VentureBeat (openai.com returned errors to our fetcher, so this is secondary): "improvements to inference and caching allowed it to reduce prices while increasing capability." Plus a 90% discount on cached reads, and rates that are "permanent prices, not promotional or introductory pricing." For GPT-6.1 Sol on 29 September the coverage gives no reason. The price stayed $2 and $10; only the cached-input rate halved, from $0.20 to $0.10.

People also reach for hardware. In April Anthropic said it runs Claude on AWS Trainium, Google TPUs and NVIDIA GPUs, and expects new multi-gigawatt TPU capacity "starting in 2027". The page links none of that to price.

QuestionAnthropicOpenAI (as reported)
Serving got cheaper?Yes for Opus 5.5: "less compute to serve"Says "inference and caching" improved
Model uses fewer tokens?Yes, says AnthropicNot claimed; Artificial Analysis found slightly more
Names a technique, kernel, chip or number format?NoNo, beyond "inference and caching"
Says why cache reads got cheaper?NoNames caching improvements, no detail

The third row is the whole problem. Nobody has said quantisation, sparse attention, speculative decoding, disaggregated serving, a chip generation or a smaller model. Anyone who says Opus 5.5 is cheaper because of one of those is guessing. So are we, openly.

Three different numbers

Opus 5.5 is 20% cheaper per token than Opus 5. Anthropic estimates roughly 40% less per task at default settings on typical workloads. That is a per-task estimate, not a per-token cut. If input and output costs moved together, the arithmetic is 0.60 ÷ 0.80 = 0.75: the remaining gap implies roughly a quarter less token volume per task on top of the rate cut. Cache reads fell 60% and your mix matters, so treat that as a shape, not a forecast.

Sonnet 5.5 is pure token behaviour, because the rate is unchanged. OpenAI is the opposite. Artificial Analysis reported that Sol and Luna used slightly more output tokens per task than GPT-5.6, so OpenAI's saving is a rate cut and nothing else.

Fewer tokens does not mean lean. On Artificial Analysis Intelligence Index v4.3.2, Opus 5.5 at max effort used about 260 million output tokens against a median of 81 million ("very verbose"). Sonnet 5.5 at max used 410 million, or $7.60 per Index task despite the low rate card.

So only one slice of "prices are falling" needs a serving-cost theory: the rate cuts. The rest is behaviour and pricing choices.

The memory problem: KV cache, and why cache prices moved

Every time a model writes a token, it consults notes on everything that came before. Those notes are the KV cache. In a long coding session the same enormous prefix is re-read again and again, which is why cache reads dominate agentic bills, as Anthropic says. Serving long context is often a memory problem rather than an arithmetic one.

Our inference: if memory is the bottleneck, making it cheaper to hold and reuse context is the most direct route to cheaper tokens, and cache-read prices are where you would see it first.

Think of a hotel that used to reserve a whole floor for every guest who might stay a month, and now allocates rooms as guests actually need them. That is PagedAttention. Kwon and colleagues reported that vLLM improved throughput "by 2-4x with the same level of latency" against FasterTransformer and Orca, through "near-zero waste in KV cache memory" (arXiv:2309.06180, 2023).

A second route is to shrink the notes. DeepSeek-V2's latent attention "reduces the KV cache by 93.3%" and lifted maximum generation throughput 5.76 times, both against DeepSeek's own 67B model (arXiv:2405.04434). A third is to read fewer notes: DeepSeek-V3.2's sparse attention "substantially reduces computational complexity", though its abstract gives no cost figures (arXiv:2512.02556).

This month's open models show these ideas in released form. Xiaomi's MiMo-V2.6-Pro, open-sourced on 22 September (China time), mixes 60 sliding-window attention layers with 10 global ones and uses 8 KV heads, per its Hugging Face card. Alibaba's Model Studio listing for DeepSeek-V4.1-Flash says "KV cache usage much lower than the prior generation" (a hosting partner's wording, not DeepSeek's).

Opus 5's cache read was $0.50; Opus 5.5's is $0.20, 0.05 times its input rate. OpenAI's GPT-6 Sol cached reads are 10% of input, and GPT-6.1 Sol's 5%. Opus 5.5 and GPT-6.1 Sol therefore land on the same ratio; that is our observation, not theirs.

Not confirmed: that any of these techniques sits inside Opus 5.5 or GPT-6. A cache-read price is also a commercial lever, because a hit skips recomputation and rewards customers for building cache-friendly prompts. The fall is consistent with cheaper memory handling. It is not proof of it.

Experts, drafters and short numbers

Mixture of experts. A hospital does not send you past every doctor. It routes you to the two specialists you need. A mixture-of-experts model has a huge total parameter count but activates a small slice per token, and cost per token tracks the slice. MiMo-V2.6-Pro has 1.02 trillion parameters, 42 billion active, and 384 routed experts with 8 active per token. DeepSeek-V3 has 671 billion total and 37 billion active, and its report says full training took "only 2.788M H800 GPU hours" (arXiv:2412.19437).

The catch is serving. Spread hundreds of experts across many GPUs and every token triggers traffic between them. DeepSeek's open-source DeepEP library exists for exactly that: "all-to-all GPU kernels for MoE dispatch and combine, including FP8 dispatch" (its page gives no benchmark numbers). Neither Anthropic nor OpenAI has disclosed its architecture, and we are not saying Opus 5.5 or GPT-6 is a mixture of experts.

Speculative decoding. A junior writes the next paragraph quickly and a senior checks it in one pass, keeping what is right and rewriting what is not. The finished text is what the senior would have written. Leviathan, Kalman and Matias reported "2X-3X acceleration" on T5-XXL "with identical outputs" (arXiv:2211.17192). MiMo-V2.6-Pro ships a five-layer multi-token-prediction drafter that predicts 7 tokens per pass. Anthropic says Opus 5.5 "generates output more than 30% faster than Opus 5" and names no reason. Our inference: hardware that finishes each request sooner can serve more requests, so speed and cost are linked.

FP4. Rounding every number in a spreadsheet to fewer decimal places makes the file smaller and quicker to move, and with careful scaling the answers barely change. FP4 is that, applied to model weights and activations. This is the thinnest part of our case. The one FP4 paper we can cite, NVIDIA's on NVFP4, trained a 12-billion-parameter model on 10 trillion tokens with results "comparable to an FP8 baseline" (arXiv:2509.25149). It is a pretraining result, not an inference-cost result, and we do not use it as one. NVIDIA's blog lists NVFP4 among the techniques in its serving stack but gives no per-technique figure. No vendor has said any frontier model is quantised to four bits, and we do not claim it.

Kernels and scheduling: the unglamorous work

The last layer is how well software keeps expensive chips busy. Picture the same recipe cooked in two kitchens, one where the ovens sit idle between dishes and one where the workflow keeps them full.

FlashAttention-3 reported a "1.5-2.0x" speedup on H100 GPUs, reaching 740 TFLOPs/s in FP16, 75% utilisation (arXiv:2407.08608). Its successor FlashAttention-4 reported up to 1.3 times cuDNN 9.13 and 2.7 times Triton on B200 GPUs at BF16, peaking at 1,613 TFLOPs/s (arXiv:2603.05451). We do not claim any named frontier model uses FlashAttention-4.

Scheduling matters as much. DistServe split the two phases of a request, reading the prompt and writing the answer, onto separate hardware, like a prep kitchen and a grill line. It reported serving "7.4x more requests or 12.6x tighter SLO" against the systems of its day (arXiv:2401.09670, 2024). The batching layer is in continuous batching explained and the vLLM deep dive; the demand side is in batch inference for cost.

NVIDIA's blog of 30 June says disaggregated serving, large expert parallelism, NVFP4 and multi-token prediction "combined, they increase throughput by up to 20x", and that token costs fell "up to 5x on the DeepSeek V4 model in just one month". Read that as NVIDIA's claim about its own stack on an open model, per GPU, with no frontier model named. It is not an industry-wide result.

TechniquePlain pictureBest public evidenceVendor-confirmed for Opus 5.5 or GPT-6?
Paged and compressed KV cacheRooms allocated as neededPagedAttention 2-4x; MLA 93.3% smaller cacheNo
Sparse and sliding-window attentionRead fewer notesDeepSeek-V3.2 abstract; MiMo-V2.6-Pro cardNo
Mixture of expertsSend the patient to two specialistsMiMo 42B of 1.02T active; DeepSeek-V3No
Speculative decodingJunior drafts, senior checks2X-3X, identical outputs; MiMo drafterNo
FP4Fewer decimal placesNVFP4 pretraining paper (not serving)No
Faster kernelsOvens never idleFlashAttention-3 and -4 abstractsNo
Disaggregated servingPrep kitchen, grill lineDistServe 7.4xNo

Every cell in the right-hand column is "No". That column is why this is a theory.

What else could explain the cuts

A cost story is not the only one that explains a price cut.

Competition. Prices follow rivals as much as serving cost. Elsewhere in September, xAI's Grok 4.7 was listed at $2 in and $6 out under 200,000 prompt tokens, and Meta's Muse Spark 1.3 at $1.25 and $4.25. Artificial Analysis scores the open-weights MiMo-V2.6-Pro 46 on Index v4.3.2 at about $0.13 per Index task. GPT-6 Sol's $2 and $10 also equals Sonnet 5.5's rate card exactly. A firm can cut price with no change in cost.

Token efficiency from training, not serving. Our inference: fewer tokens per task comes from how a model is trained to reason and use tools. Anthropic's customer quotes, which are vendor-curated, describe Box using about a third of Opus 5's tokens and Factory 20-25% fewer output tokens. That is model behaviour, and it could arrive with no serving change at all.

Cache pricing as strategy. A steep cache discount nudges customers toward patterns cheaper to serve. It can be generous without the cost having fallen by the same amount.

Capacity, mix and margin. We have no data on any of them. Anthropic's new TPU capacity is not due until 2027, so it cannot have caused a September cut, and we credit no chip generation.

Speculation about the model itself. Some third-party commentators have suggested Opus 5.5 is a smaller or distilled model, or runs quantised or on re-architected serving. That is unconfirmed opinion, and we make no claim either way.

What would prove us wrong

A theory must be able to lose. Six things would move us.

  1. A vendor names a different cause. A system card or engineering post crediting capacity, margin, a smaller model or a business decision would displace the serving-technique story.
  2. A cut is reversed. OpenAI says its rates are permanent. If either company walks a cut back, the story is pricing strategy, not cost.
  3. The cut does not spread down the range. Sonnet 5.5's rate did not move. Anthropic says Haiku 5.5 will "follow in the coming weeks", which is a live test. We make no prediction on its price.
  4. Matched-token comparisons erase the saving. If cost per task, with token counts held equal, shows no difference from the older models, the story is behaviour, not serving.
  5. Independent throughput does not move. Anthropic's "more than 30% faster" is a vendor claim. If independent measurement at the same price shows no gain, the serving story loses its best supporting fact.
  6. Hardware timing. If the next large cuts arrive only when the 2027 TPU capacity does, hardware deserves more credit than we give it.

Our own scorecard, honestly. That serving got cheaper is partly confirmed, by Anthropic's words and by OpenAI's as reported. That the cause is the technique family above is our inference: plausible, well cited, and not confirmed for either company. That price followed cost is the least certain link of all.

Where Swfte sits

We do not run Anthropic's or OpenAI's serving stacks, and nothing here comes from inside them. It is a reading of public sources.

What we do offer is narrow. Swfte Connect is "One API for Every AI Model": "Connect to 50+ LLM providers through a single, unified API. Smart routing optimizes cost, latency, and availability." If you cannot predict which vendor cuts next, a single API in front of many providers keeps you ready. Swfte Studio records "traces, logs, token and cost telemetry", which is what you need to measure cost per completed task rather than cost per token. The token cost calculator and AI model leaderboard are the quickest way to test the arithmetic above on your own numbers.

What to do if you pay the bill

If you are the CFO. Budget on cost per completed task, not the rate card. A hypothetical: a team spends $10,000 a month on Opus 5 tokens. A pure 20% rate cut takes that to $8,000. Anthropic's 40% per-task estimate would take it to $6,000, but only if your workload behaves like its typical one. Test your own number before banking it, and read where token prices are going before committing multi-year spend to continuing declines.

If you are the CTO. Measure tokens per task, cache-hit share of input, and effort setting. Cache reads are the lever Anthropic itself points to, so keep prompt prefixes stable. A hypothetical: 1 billion cached tokens a month cost $500 at Opus 5's $0.50 and $200 at Opus 5.5's $0.20. The reasoning-token piece explains why effort settings can swamp a rate card.

If you run your own inference. The techniques are public and open models already use several. The DeepSeek V4.1-Flash self-hosting guide is a practical start, and the open-source cost guide covers when it pays.

If you are choosing between vendors. A cheaper rate is not a cheaper task. Route by workload, as in intelligent LLM routing, and re-run the comparison whenever a vendor moves.

Prices fell in September, and the plausible reasons are sitting in public papers. What actually changed inside Anthropic's and OpenAI's serving systems is still known only to them, and the sensible position is to enjoy the lower bill without building a forecast on a mechanism nobody has confirmed.


Related: Claude Opus 5.5 and Sonnet 5.5, GPT-6 Sol, Luna and 6.1 Sol, The Efficiency Race, current rankings on AI model leaderboard.

Sources: Anthropic, Claude Opus 5.5 · Anthropic, Claude pricing · Anthropic, Google and Broadcom compute · VentureBeat, GPT-6 Sol and Luna · VentureBeat, GPT-6.1 Sol · Xiaomi, MiMo-V2.6-Pro card · Alibaba Model Studio, newly released models · NVIDIA, inference software and token cost · DeepSeek, DeepEP · Meta, Model API pricing · xAI release notes

Prices and quotes were checked against vendor pages as recorded in our fact sheet on 30 September 2026. Anthropic and OpenAI benchmark and cost figures are vendor-run and not independently reproduced; OpenAI statements are as reported by VentureBeat. Artificial Analysis figures are Intelligence Index v4.3.2. The papers are cited from their abstracts.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Deploy a model with Swfte Connect

One gateway, every provider, per-token cost visibility. Swap models without touching your code.