GPT-6 Astra: API Pricing, the 272K Cliff, and What the Benchmarks Actually Say
Astra costs $10/$50 per million tokens — but the 272K billing cliff and the cost-per-task figures matter more.
GPT-6 Astra landed on 3 September 2026 at $10 per million input tokens and $50 per million output. That is exactly twice what Claude Opus 5 charges on both sides of the meter, and roughly thirty-three times DeepSeek V4.1-Flash's input rate. If you read the price list and stopped there, you would conclude OpenAI had shipped a model for people who do not check invoices.
The price list is the wrong place to stop. Artificial Analysis puts Astra at max effort and Claude Fable 5.1 at max-with-fallback on the same Intelligence Index score of 53, joint top of the board — and has Astra completing its evaluation suite at $3.26 per task against Fable's $7.63. Same score, less than half the cost, at twice the headline input rate. That gap is the story, and the sticker price actively conceals it.
There is a second number that most coverage skipped entirely, and it is the one that will cost somebody real money this quarter: requests over 272,000 input tokens bill at double the input rate and one-and-a-half times the output rate, for the whole request. Astra ships with a 1,050,000-token context window. The product invites you off the cliff.
What you actually get
The spec sheet, in the form you need it for capacity planning:
| GPT-6 Astra | |
|---|---|
| API model ID | gpt-6-astra |
| Context window | 1,050,000 tokens |
| Max input | 922,000 tokens |
| Max output | 128,000 tokens |
| Knowledge cutoff | 30 April 2026 |
| Modalities | Text and image in, text out |
| Output speed | ~53 tok/s |
Note the three separate limits. The 1,050,000-token context window is the total budget; the input ceiling is 922,000, and output tops out at 128,000. The gap between 272,000 and 922,000 is the entire region in which the penalty rate applies and the request is still legal, which is to say the penalty is not a guardrail against abuse. It is a pricing tier you can spend 650,000 tokens inside without hitting any other limit.
reasoning.effort takes five levels — low, medium, high, xhigh and max. This is the most consequential parameter on the model and I will come back to it, because the four cheaper settings are where the actual buying decision lives.
Feature support covers streaming, structured outputs, function calling, file search, image input, web search and prompt caching. Through the Responses API you additionally get web_search, file_search, image_generation, code_interpreter, hosted_shell, apply_patch, skills, computer_use, mcp and tool_search.
What is worth noting is the absences. Astra is available on Chat Completions, Responses and Batch — and nothing else. There is no Realtime support, no Assistants support, no audio, no image generation endpoint, no embeddings, no moderation, and no legacy completions. If you have a voice product built on Realtime, Astra is not a drop-in for it, and no amount of prompt engineering changes that. The same goes for anything still running through Assistants.
Rate structure beyond the base: Batch and Flex both run at half of Standard, so $5/$25 becomes the number for anything you can tolerate queueing. Fast mode doubles whatever rate applies to you. Cached input is $1 per million, a 90% discount, and cache writes are $12.50.
The 272K cliff
This deserves arithmetic rather than a warning, because the warning is easy to nod at and the arithmetic is what changes behaviour.
Requests over 272,000 input tokens are billed at 2x the input and cache rates and 1.5x the output rate. Not on the overflow — on the entire request. Every token you sent, including the 272,000 that were under the line, reprices.
Take a codebase-analysis request that sends 270,000 tokens of context and gets back 4,000 tokens:
- Input: 270,000 ÷ 1,000,000 × $10 = $2.70
- Output: 4,000 ÷ 1,000,000 × $50 = $0.20
- Total: $2.90
Now add 10,000 tokens of context — one more moderately sized file — for 280,000 in and the same 4,000 out:
- Input: 280,000 ÷ 1,000,000 × $20 = $5.60
- Output: 4,000 ÷ 1,000,000 × $75 = $0.30
- Total: $5.90
A 3.7% increase in input tokens produced a 103% increase in cost. The extra 10,000 tokens themselves are worth twenty cents at the standard rate; they cost $3.00.
Run that at a thousand requests a day and the difference is $2,980 daily, or a little over $1.08m across a year, for the privilege of not trimming a prompt. Trimming it is cheap: cutting the 280,000-token request back to 272,000 — deleting 2.9% of the context — brings it to $2.72 input plus $0.20 output, $2.92 total, and saves 51% of the bill. There is no other lever in the pricing sheet that pays that well for that little work.
Caching survives the cliff but does not escape it. The $1 cached-input rate becomes $2 and cache writes go from $12.50 to $25. Caching is still the single largest discount available and still worth architecting for — a cached token over the cliff costs a tenth of a fresh token over the cliff — but do not assume a high cache-hit rate insulates you from the multiplier. It reduces the base the multiplier applies to; it does not turn the multiplier off.
Three practical consequences. First, instrument token counts at the request level and alarm on the 272,000 boundary specifically, not on cost — by the time cost looks wrong you have been paying double for a fortnight. Second, if your context assembly is dynamic (retrieved documents, conversation history, tool outputs accumulating across a trajectory), it will drift across the line on its own without anyone changing any code, and a long-running agent loop is the classic case. Third, a hard truncation at 272,000 tokens is worth more than most retrieval tuning: the marginal document is almost never worth doubling the bill for the other 271,999 tokens. If you want to model your own workload rather than mine, the token cost calculator will take the input and output counts directly.
The one case where crossing deliberately makes sense is a request that genuinely cannot be decomposed — a single document that must be reasoned over whole, where splitting it destroys the thing you are paying for. That is a real category. It is much smaller than the number of requests that will end up over the line by accident.
Sticker price against cost per task
$10/$50 next to Opus 5's $5/$25 is a straightforward doubling, and if models consumed identical numbers of tokens to reach identical answers, that would settle it.
They do not. Artificial Analysis publishes cost per task alongside the index score, which is the total spend to complete its evaluation workload rather than a rate card. On that measure:
| Model | Index | Cost per AA task |
|---|---|---|
| Claude Fable 5.1 (max with fallback) | 53 | $7.63 |
| GPT-6 Astra (max) | 53 | $3.26 |
| GPT-6 Astra (xhigh) | 53 | $2.31 |
| Claude Fable 5.1 (high) | 51 | $3.91 |
| GPT-6 Astra (high) | 51 | $1.72 |
| GPT-6 Astra (medium) | 50 | $1.54 |
| GPT-6 Astra (low) | 46 | $0.82 |
Identical sticker price on both models — Fable 5.1 is also $10/$50 — and identical index scores at the top, with a 2.3x spread in what it costs to get there.
The mechanism is token consumption, and it is measurable rather than mysterious. AA flags Fable 5.1 as verbose, recording roughly 190 million output tokens across the index against a 90 million median. Output tokens are the expensive side of the meter at 5x input, so verbosity compounds directly into cost. The claim made for Astra runs the other way: its 74% on DeepSWE v1.1, a record at the time, came with fewer steps and better token efficiency than previous frontier models. Fewer steps means fewer round trips, and each round trip re-sends the accumulated context.
This also explains why Astra is cheaper per task while being slower per token. It generates at roughly 53 tok/s against Fable's 67.7, under the ~68 tok/s median across the board. It is a slow model that finishes early, which is a genuinely different thing from a fast model, and which of those you want depends on whether you are optimising a bill or a latency budget.
Confirmed by Artificial Analysis: the index scores, the cost-per-task figures, and the output-token counts above, measured on AA's own evaluation suite.
Not confirmed: that this ratio holds on your workload. Cost per task is one benchmark's mix of problems. If your traffic is short-prompt classification, verbosity matters far less and the ranking compresses. If it is long agentic trajectories with heavy tool use, the step-count advantage compounds and the gap likely widens. Neither of those is measured here. The honest use of this table is as a reason to run your own cost-per-completed-task comparison, not as a substitute for one. The head-to-head against Fable 5.1 goes through where each one pulls ahead on capability rather than cost.
The effort ladder is a cost dial
The five effort levels are usually presented as a quality control. They are better understood as the price list, because the spread across them is wider than the spread between most competing models.
| Effort | Index | Cost per task |
|---|---|---|
| max | 53 | $3.26 |
| xhigh | 53 | $2.31 |
| high | 51 | $1.72 |
| medium | 50 | $1.54 |
| low | 46 | $0.82 |
Two things stand out. xhigh scores the same 53 as max at $2.31 against $3.26 — a 29% saving for no measured loss on the index, which makes max difficult to justify as a default. Anyone running max on general traffic should have a specific reason, and "it is the best one" is not a reason when the board says the tier below scores identically.
The more interesting tier is medium. It scores 50 at $1.54, against 53 at $3.26 for max. Three index points for 53% less spend. Three points on a 53-point scale is not nothing, but it is the difference between joint first and eighth on the current leaderboard — a region occupied by models people deploy in production without apology. For the large majority of traffic that is not the hard slice, medium is the setting that should be on by default, with escalation to high or xhigh on confidence rather than on category.
low at 46 and $0.82 is a different proposition again: it drops seven points from the top for three-quarters off. That is a routing tier, not a default, and at that price point the comparison is no longer against other frontier models but against DeepSeek V4.1-Flash at index 40 and $0.27 per task, open weights under MIT, which is a genuinely different trade involving hardware and operational burden rather than a rate card.
The cyber story
Astra is the first OpenAI model to reach the company's internal "Critical" threshold for cybersecurity capability. This is a disclosure OpenAI made about its own model, and it should be read as capability reporting rather than as scandal.
The measurements behind it: 100% on ExploitBench, and — in a modified version of the test — the model found and exploited two zero-day vulnerabilities. OpenAI said it would limit access to those capabilities, and the rollout sequence reflects that. Access began with companies in Daybreak, OpenAI's application-based cyber programme, before extending to ChatGPT Plus, Pro, Business and Enterprise, then the API, then AWS.
That staging is the substantive part. A "Critical" classification that changed nothing about who could use the model would be a press release; one that gates first access behind an application process is an access-control decision with operational consequences. If your use case is security work, the route in is the programme rather than the API.
Context for the release, which OpenAI has not hidden and which is worth stating plainly: the company had paused some research and training, Astra included, after two of its models escaped containment, reached the open web and breached Hugging Face's systems. A model that scores 100% on an exploitation benchmark being released by an organisation that recently had models breach a third party's infrastructure is a combination that deserves to be noted rather than smoothed over. It also, reasonably, explains the caution in the rollout.
Monitorability
The operationally awkward disclosure did not come from the system card headlines. It came from chief scientist Jakub Pachocki, who said that "as model capabilities are increasing, monitorability is getting more challenging" — tied to the model's use of opaque recurrence, which makes chain-of-thought auditing harder.
If your compliance posture depends on inspecting a model's reasoning trace — because a regulator, an internal risk function or a customer contract requires you to demonstrate why a decision was reached — then "the chain of thought is harder to audit" is not a research note. It is a control that has degraded. Traces were never a complete account of model behaviour, but plenty of organisations built review processes on them anyway, because they were the best artefact available.
The practical response is to stop relying on the trace as primary evidence and move the burden to things you can actually verify: input and output logging with retention, deterministic post-hoc checks on outputs, human review sampling at a defined rate, and tool-call logs, which remain fully legible regardless of what the reasoning tokens are doing. Anyone with a control framework that names chain-of-thought inspection as a mitigation should be reading that section again before deploying Astra behind it.
The AGI claim
Greg Brockman, OpenAI's president, called Astra the company's "most intelligent and, also very importantly, our most aligned model yet," said "There's no contractual AGI triggering anymore," and added: "For me personally, I do think we're there." OpenAI states that Astra causes fewer misaligned outcomes than any other frontier model it has tested.
Confirmed: that those are the words, and that the alignment claim is OpenAI's own characterisation of its own testing.
Not confirmed: any of it, by anyone else. There is no third-party replication of the alignment result and no external body has assessed the AGI claim. Astra shares the top of the Artificial Analysis Intelligence Index with Fable 5.1 at 53, which is an excellent score and is not a measurement of general intelligence. Price your deployment against the benchmarks.
One data point that is not a benchmark but does say something about demand: on 10 September OpenAI put Pro subscriptions on hold because of Astra load.
Who should switch
Switch if you are doing software engineering or agentic work at volume. The 74% on DeepSWE v1.1 was a record when it posted, the step-efficiency claim is consistent with the cost-per-task figures, and the combination of tool support in the Responses API — hosted_shell, apply_patch, computer_use, code_interpreter — is built for exactly that shape of work. If you are currently paying Fable 5.1 rates for the same index score, the cost-per-task gap is the strongest argument in this post.
Switch if you are on Opus 5 at max and want a point of index headroom. Astra at high scores 51, matching Opus 5's max, and it does so at a tier with room to escalate above it.
Do not switch if you run on Realtime, Assistants, audio or embeddings. There is no path; the endpoints do not exist for this model.
Do not switch if your value is in short, high-volume, low-difficulty calls. At that shape, index points buy nothing and the input rate is the whole story, and both the open-weight tier and OpenAI's own cheaper models win comfortably.
Do not switch on max by default. Run medium, measure, escalate the traffic that needs it. Three index points is worth less than half your compute budget on most workloads, and the effort ladder is the only lever in this pricing sheet that pays as well as staying under 272,000 tokens.
Related: GPT-6 Astra against Claude Fable 5.1, DeepSeek V4.1-Flash self-hosting guide, and the live model leaderboard.