GPT-6 Sol Pricing After the Cut: Luna, 6.1 Sol and the Opus 5.5 Question
OpenAI halved GPT-6 Sol and Luna prices and says the new rates are permanent. What it means for your default model.
Picture a product lead who locked in a default model on 21 September. A frontier Sol-class model at what VentureBeat's arithmetic implies was $4 per million input tokens and $20 per million output, an eval suite passed and a budget signed. On 22 September OpenAI released GPT-6 Sol and GPT-6 Luna at half the price of the previous generation. On 29 September it released GPT-6.1 Sol at the same $2/$10 as GPT-6 Sol, with a higher score and a cached-input rate cut in half again. That product lead is hypothetical. The problem is not.
OpenAI's docs now list three models at $2/$10, $0.10/$0.50 and $2/$10. An OpenAI spokesperson, as reported by VentureBeat, says the Sol and Luna rates are permanent prices and not promotional ones. If you believe that, a price you can plan around has arrived at the top end of the market, and the interesting question moves from "what does it cost per token" to "what does it cost to finish a job".
That is where this post spends its time: the docs, the 272K surcharge that undoes part of the saving, what the cut reportedly came from, Artificial Analysis scores by effort setting, and a price-per-task comparison against Claude Opus 5.5 and Sonnet 5.5. Arithmetic is shown wherever I derive something, and vendor-run or single-outlet claims are labelled.
What OpenAI's docs list
These figures come from OpenAI's developer documentation pages for each model.
| GPT-6 Sol | GPT-6 Luna | GPT-6.1 Sol | |
|---|---|---|---|
| API ID | gpt-6-sol | gpt-6-luna | gpt-6.1-sol |
| Input / output per 1M tokens | $2 / $10 | $0.10 / $0.50 | $2 / $10 |
| Cached input | $0.20 | $0.01 | $0.10 |
| Cache write | $2.50 | $0.125 | $2.50 |
| Context / max input / max output | 1,050,000 / 922,000 / 128,000 | same | same |
| Knowledge cutoff | 20 April 2026 | 18 May 2026 | 30 April 2026 |
| Released (as reported) | 22 September | 22 September | 29 September |
The release dates are not on the docs pages; they come from VentureBeat and TechCrunch, and Artificial Analysis lists the same three dates.
Three details are worth pulling out.
The limits are the same across the family. All three carry a 1,050,000-token context window, a 922,000-token input ceiling and 128,000 tokens of output. Artificial Analysis lists GPT-6 Sol at 872K. OpenAI's docs say 1,050,000, so that is the vendor spec I use; if you cite AA's page, say "AA lists 872K".
The 6.1 Sol change is in the cache. Standard input and output are identical to GPT-6 Sol. The one line that moved is cached input, from $0.20 to $0.10, which is 5% of the input rate rather than 10%. For agent loops that resend a long, stable prefix on every turn, that is the line that matters most.
Modifiers stack on top. OpenAI lists Batch and Flex at 50% off, Fast mode at 2x, and regional processing at +10%. For 6.1 Sol the docs show effort options of low, medium (the default), high, xhigh and max, with text and image in, text out. I did not re-check those for Sol and Luna.
For scale, the model OpenAI has been positioning against is GPT-6 Astra at $10/$50, which we covered in the GPT-6 Astra pricing breakdown. VentureBeat says 6.1 Sol's standard input and output are exactly one-fifth of Astra's. On the rate card, that is right: $2 against $10, and $10 against $50.
Halved, permanent, and the reason as reported
VentureBeat quotes OpenAI saying that "improvements to inference and caching allowed it to reduce prices while increasing capability." I could not read OpenAI's own announcement page, which returned an access error, so treat that quote as reported rather than primary. The same coverage says GPT-6 carries a 90% discount on cached input-token reads, which matches the docs: $0.20 against $2 for Sol, $0.01 against $0.10 for Luna.
VentureBeat calls Sol "exactly 50% cheaper in both directions" than GPT-5.6 Sol, which implies $4/$20 for the predecessor, and puts Luna at 50% cheaper on input and 58.3% cheaper on output. Those predecessor prices are the outlet's arithmetic, not something I verified on an OpenAI page. One earlier lead said the comparison was against promotional pricing. It came from a low-quality aggregator, and I have not repeated it.
What is worth noticing is what the saving is made of. Artificial Analysis found that the Sol and Luna savings were "driven by price cut", with the models using slightly more output tokens per task, about 31K against 29K for Sol and 51K against 41K for Luna. It put Sol's cost per Intelligence Index task at $1.06 against $1.99 for GPT-5.6 Sol. Do the sum: a 50% rate cut multiplied by roughly 7% more tokens (31 ÷ 29) gives about 0.53 of the old bill, a fall of around 47%. The measured figure, $1.06 ÷ $1.99, is 0.53. It fits.
So OpenAI's cheaper was a cheaper rate, not a more economical model. That is the opposite shape from Anthropic's Opus 5.5 story, where the rate fell 20% and tokens per task also fell, which Anthropic estimates at about 40% less per task than Opus 5. Our companion post on the two Claude 5.5 releases goes through that in full. I have a working theory about why prices are falling across vendors, with the papers behind it and a clear list of what is not confirmed, in why frontier AI got cheaper. The short version is that OpenAI has named "inference and caching", Anthropic has named no mechanism, and nobody has named a chip or a kernel.
For 6.1 Sol, OpenAI's coverage gives no reason for the price at all. It is the same as Sol's, and the thing that changed is capability and the cache rate. OpenAI says the model "nearly matches GPT-6 Astra's intelligence" on agentic coding, computer use and professional work, at one-fifth of Astra's standard token prices, per TechCrunch quoting OpenAI.
The 272K surcharge, with numbers
The docs for all three models carry a rule that the headline rate does not mention. Requests over 272K input tokens are billed at 2x for input and cache, and 1.5x for output, across the whole request, not just the overflow. It is the same cliff we flagged on Astra, and it sits inside a 922K input ceiling, so you can spend 650K tokens in the penalty zone without hitting any other limit.
Take a GPT-6.1 Sol request with 270,000 input tokens and 4,000 output tokens:
- Input: 0.27 × $2 = $0.54
- Output: 0.004 × $10 = $0.04
- Total: $0.58
Add 10,000 tokens of context, so 280,000 in and the same 4,000 out:
- Input: 0.28 × $4 = $1.12
- Output: 0.004 × $15 = $0.06
- Total: $1.18
That is 3.7% more input and a 103% higher bill. Trimming back to 272,000 tokens costs 0.272 × $2 + $0.04 = $0.584, about half. Cached reads cross the line too: $0.10 becomes $0.20.
Now the part that matters for the model choice. Anthropic's pricing page says Claude 4.6 and later models include the full 1M-token context at standard pricing, so Opus 5.5 and Sonnet 5.5 have no long-context surcharge. Run the same 280K-in, 4K-out request through them:
- Opus 5.5 at $4/$20: 0.28 × $4 + 0.004 × $20 = $1.12 + $0.08 = $1.20
- Sonnet 5.5 at $2/$10: 0.28 × $2 + 0.004 × $10 = $0.56 + $0.04 = $0.60
- GPT-6.1 Sol: $1.18
Past the line, 6.1 Sol's list-price advantage over Opus 5.5 disappears, and Sonnet 5.5 costs half as much as either. This assumes identical token counts, which will not hold in practice. But if your workload sits routinely at 300K or 600K tokens of context, the "OpenAI halved its prices" story does not apply to you, and your cheapest option on the rate card is somebody else's.
Price per token is not price per task
The rate card tells you what a token costs. It does not tell you how many tokens a model uses to finish something. Artificial Analysis publishes both a score and a cost per Intelligence Index task, and everything below is from the v4.3.x index, with the effort setting named on each row.
| Model (effort) | AA Index v4.3.x | Cost per Index task |
|---|---|---|
| GPT-6.1 Sol (max) | 52 | $0.72 |
| GPT-6.1 Sol (xhigh) | 51 | $0.39 |
| GPT-6.1 Sol (high) | 50 | $0.32 |
| GPT-6.1 Sol (medium) | 48 | $0.21 |
| GPT-6 Sol (max) | 48 | $1.05 |
| GPT-6 Sol (xhigh) | 44 | $0.52 |
| GPT-6 Luna (max) | 37 | $0.07 |
| GPT-6 Astra (max) | 53 | $3.26 |
Four things come out of it.
6.1 Sol is the better Sol at the same price. At max effort it scores 52 against 48 for GPT-6 Sol, at $0.72 against $1.05 per task, on an identical rate card. AA lists 67M output tokens for 6.1 Sol against 77M for GPT-6 Sol, so part of the drop is simply fewer tokens. GPT-6 Sol at max, 48 for $1.05, is matched by 6.1 Sol at medium, 48 for $0.21. Same score, one-fifth the cost.
Against Astra, 6.1 Sol is one point behind at 4.5x less per task. Astra at max is 53 for $3.26. 6.1 Sol at max is 52 for $0.72. $3.26 ÷ $0.72 is 4.5. This is consistent with OpenAI's "nearly matches", and it is an independent measurement rather than a vendor one.
The effort dial is the cheap lever. Going from max to medium on 6.1 Sol costs four index points (52 to 48) and saves 71% ($0.72 to $0.21). From max to high costs two points and saves 56% ($0.72 to $0.32). Medium is the default in OpenAI's docs, which is a sensible default.
Luna shows why per-token is misleading. Luna's rate card is 20x cheaper than Sol's ($0.10 against $2 on input). Per Index task, Luna at max is $0.07 against $1.05 for GPT-6 Sol at max, which is 15x cheaper, not 20x. The reason is verbosity: AA lists 140M output tokens for Luna against 77M for Sol, about 1.8x. Luna is also 11 points lower on the index (37 against 48). It is a bulk-work model, and I would route to it deliberately rather than default to it.
One caveat governs the whole table. AA's cost per task is one benchmark's mix of problems. Your traffic may be shorter, longer, more tool-heavy or more cache-friendly. Run your own cost-per-completed-task test, and use the token cost calculator to turn token counts into money. Our post on hidden reasoning-token costs explains why.
Head to head with Opus 5.5 and Sonnet 5.5
Now the comparison a product team actually faces. Rate cards first, from OpenAI's and Anthropic's own pages:
| Input / output | Cache hit | Context | Long-context surcharge | |
|---|---|---|---|---|
| GPT-6.1 Sol | $2 / $10 | $0.10 | 1,050,000 | Yes, over 272K |
| Claude Opus 5.5 | $4 / $20 | $0.20 | 1M | None |
| Claude Sonnet 5.5 | $2 / $10 | $0.20 | 1M | None |
Per token, Sonnet 5.5 and 6.1 Sol are identical, and Opus 5.5 is exactly double. Cached reads are half the price on 6.1 Sol. Then the per-task view from AA v4.3.x:
| Model (effort) | AA Index v4.3.x | Cost per Index task |
|---|---|---|
| Claude Opus 5.5 (max) | 58 | $5.98 |
| Claude Opus 5.5 (xhigh) | 56 | $3.46 |
| Claude Opus 5.5 (high) | 54 | $1.82 |
| Claude Opus 5.5 (medium) | 51 | $1.34 |
| Claude Sonnet 5.5 (max) | 56 | $7.60 |
| Claude Sonnet 5.5 (xhigh) | 52 | $2.74 |
| Claude Sonnet 5.5 (high) | 47 | $1.08 |
| GPT-6.1 Sol (max) | 52 | $0.72 |
| GPT-6.1 Sol (medium) | 48 | $0.21 |
What the numbers say, without rounding in anyone's favour:
On capability, Opus 5.5 leads. 58 at max against 52 for 6.1 Sol at max, six points. At high effort, Opus is 54 against 50. That is a clear gap on this index.
On cost, 6.1 Sol wins by a large multiple. Opus 5.5 at max costs 8.3x as much per task ($5.98 ÷ $0.72). Some of that is the 2x rate: AA lists 260M output tokens for Opus against 67M for 6.1 Sol, about 3.9x, and 3.9 × 2 is roughly 7.8. Cache and input mix probably account for the rest; that is my inference.
At default settings the gap is about 6x. Anthropic's docs put Opus 5.5's default effort at medium, and OpenAI's default for 6.1 Sol is also medium. Opus at medium is 51 for $1.34. 6.1 Sol at medium is 48 for $0.21. Three index points cost 6.4x more ($1.34 ÷ $0.21). Sonnet 5.5's default is high, at 47 for $1.08, which is one point lower than 6.1 Sol's medium for 5.1x the cost.
Sonnet 5.5 at max is the trap. It scores 56, close to Opus, but at $7.60 per task, more than Opus at max, because AA lists 410M output tokens for it, against a median of 81M. A cheaper rate did not make a cheaper task. Sonnet's better rows are high and xhigh.
Matched scores favour OpenAI. Sonnet 5.5 at xhigh scores 52, same as 6.1 Sol at max, at $2.74 against $0.72. That is 3.8x.
Both vendors also publish their own comparisons, and they point in opposite directions. OpenAI says, as reported, that 6.1 Sol at medium effort is 2.2 points ahead of Opus 5.5 on AutomationBench at about one-third the cost, and that it beats Opus 5.5 on GDP.pdf at under half the cost per task. Anthropic says Opus 5.5 at default effort beats GPT-6 Astra on FrontierCode at roughly 20% of the cost per task, and that Sonnet 5.5 at High effort matches GPT-6 Sol's best FrontierCode score for about a fifth of the cost. All of these are vendor-run and vendor-selected. OpenAI's launch-day comparisons on 22 September were against Opus 5, not Opus 5.5. I do not treat any of them as independent.
Not confirmed: LMArena. Arena's text board, last updated 25 September, ranks Opus 5.5 (high) first at 1509 with a margin of plus or minus 12 from 2,307 votes, within a few points of the next three, so it is not decisive. I could not see 6.1 Sol, Luna or Sonnet 5.5 on that board, so I have no Arena comparison for them.
Ultrafast, and the Astra that reportedly is not coming
Two things from the 29 September coverage deserve a paragraph each, with their caveats.
Ultrafast. VentureBeat and other outlets report a new tier that runs up to about 300 tokens per second, billed at 6x the model's standard rate (3x Fast mode), with claims of up to 8x faster in Codex and 6x in the API. It launched for GPT-6 Astra. For GPT-6.1 Sol it was described as arriving "in the coming days" and was not live as of 29 September, so do not build a plan around it until you see it in the docs. I am also not printing dollar figures. The per-token Ultrafast prices in circulation are the outlet's own multiplication of the standard rate, not line items I could find published by OpenAI. If you need the number, take the multiplier from OpenAI's statement and do the sum on the rate that applies to you, then check it against your invoice.
GPT-6.1 Astra. The Wall Street Journal reported on 28 September that OpenAI cancelled GPT-6.1 Astra over deception and safety regressions, with Engadget, Gizmodo and The Hacker News corroborating. I could not retrieve an OpenAI statement, so this is reported, not confirmed. The same coverage says GPT-6 Astra, released 3 September, remains available. If the reporting is right, 6.1 Sol becomes the top of the 6.1 line by default, and any plan that assumed a cheaper 6.1 Astra should be revisited.
Where Swfte sits
We route between models for a living, so we have a stake in the answer to "which default". Swfte Connect is one API for 50+ LLM providers, with smart routing that optimises cost, latency and availability, and you can bring your own keys from any provider or use pooled access for simplified billing. If the table above is your reality, with a cheap tier for bulk, a mid tier for the default and a costly one for the hard slice, that routing is the mechanism. To test a change before it hits users, Swfte Studio includes testing and evaluation (golden sets, A/B) and token and cost telemetry. Our piece on intelligent LLM routing is the longer version.
What to do this week
If you are choosing a default and your prompts are under 272K tokens. Try GPT-6.1 Sol at medium effort first. On AA's index it is 48 for $0.21, and the docs' default is the same setting. Measure it on your own traffic against your current model before you trust my table. Keep the cached-input rate in the model: at $0.10 it is the cheapest on the sheet apart from Luna.
If you hit a hard slice that the default misses. Escalate to Opus 5.5 at high (54 for $1.82) rather than max (58 for $5.98). Four more points cost 3.3x more. Whether that is worth it depends on what a failed task costs you. The gap is real, and it is also the most expensive bit of the ladder.
If you routinely send more than 272K tokens. Price it with the multipliers. At 280K in and 4K out the three models land at $1.18, $1.20 and $0.60 in the example above. Sonnet 5.5 or Opus 5.5 avoid the cliff. Or truncate to 272K, which halves the 6.1 Sol bill in that example.
If most of your traffic is bulk classification, extraction or triage. Look at Luna, with its 37 on the index and its verbosity, and measure output tokens. Its $0.07 per task at max is real, and so is the 11-point gap.
If you were about to renegotiate or commit to volume. OpenAI says these prices are permanent. Vendors can still change a price. Commit to a portability layer, not a price. Run a quarterly bake-off, keep your prompts model-agnostic, and keep at least two vendors warm.
The rate cut matters. The routing decision it makes possible matters more.
Related: GPT-6 Astra pricing and the 272K cliff, GPT-6 Astra against Claude Fable 5.1, Claude Opus 5.5 and Sonnet 5.5, why frontier AI got cheaper, current rankings on AI model leaderboard.
Sources: OpenAI, GPT-6 Sol · OpenAI, GPT-6 Luna · OpenAI, GPT-6.1 Sol · VentureBeat, GPT-6 Sol and Luna · VentureBeat, GPT-6.1 Sol · TechCrunch, GPT-6.1 Sol · Anthropic, pricing · Anthropic, Claude Opus 5.5 · Anthropic, Claude Sonnet 5.5
Method note: Artificial Analysis figures are Intelligence Index v4.3.x as printed on each AA model page on 30 September 2026, with the effort setting named per row; AA cost per task is per Index task and depends on tokens used. OpenAI and Anthropic benchmark claims are vendor-run and not independently reproduced. Statements attributed to OpenAI beyond its developer docs are as reported by VentureBeat, TechCrunch and others, because openai.com pages were not retrievable.