Unico Connect
OpenAI GPT-6 Astra pricing, specifications and API changes explained for software teams
Back to Blog
AIUpdated September 26, 202618 min read

OpenAI GPT-6 Astra Explained, What Shipped and What Breaks in Your API

Vasim Gujrati

Vasim Gujrati

Solutions Architect, AI & Platforms, Unico Connect

In this article

Updated 21 September 2026. This page was first published on 11 August 2026, when Astra had been announced but not yet released. OpenAI shipped it on 3 September 2026, so we rewrote the page around the model that launched.

GPT-6 Astra went live on 3 September 2026, and OpenAI documents its price, specification and breaking changes across five separate pages. Astra removes API parameters, moves tool calling to a different endpoint, returns an error for a reasoning setting that was valid on the previous model, and applies a price surcharge to the whole request once your input crosses a threshold at 26 percent of the context window. We have pulled all of it onto one page, with the verified specifications, the full pricing table and a worked example, every breaking change, the shutdown calendar, the compliance limits that decide whether you can use the Agents API at all, and the benchmarks where Astra loses. Every figure is primary sourced from OpenAI and dated.

Quick Answer

GPT-6 Astra is the OpenAI flagship model released on 3 September 2026, available under the API model ID gpt-6-astra. It has a 1,050,000 token context window, accepts 922,000 input tokens, returns up to 128,000 output tokens, and has a knowledge cutoff of 30 April 2026. It takes text and image input and returns text only. API pricing starts at 10 US dollars per million input tokens and 50 US dollars per million output tokens, doubling on input and rising 50 percent on output for any request over 272,000 input tokens. OpenAI classified it Critical for cybersecurity under its Preparedness Framework, the first model it has ever placed at that level, and shipped it anyway with added safeguards. Unico Connect builds agentic AI systems behind model abstractions, which keeps a release like this down to a config change.

Key Takeaways

  • Astra shipped on 3 September 2026, and gpt-6-astra is the only Astra model listed in the API, so that is the ID to point your code at. No gpt-6-astra-pro, -mini or -codex model ID is published, although OpenAI does name a GPT-6 Astra Pro inside ChatGPT, which its help center calls GPT-6 Pro. A plan that depends on one of those model IDs has nothing to call, but the Responses API does offer a pro reasoning mode on GPT-6 models through reasoning.mode.
  • The pricing cliff sits at 272,000 input tokens, 26 percent of the way into the 1,050,000 token window, and crossing it prices the entire request at the higher rate, including the tokens below it. If your requests run close to the cliff, a retrieval step that keeps them under it pays for itself.
  • Move every agent that calls tools to the Responses API. Chat Completions still serves Astra for plain text, but it no longer supports tool calling.
  • Code written against GPT-5.5 breaks on Astra until you delete temperature, top_p and top_logprobs and stop sending reasoning effort none, which now returns HTTP 400.
  • For regulated clients, settle compliance first. The Agents API offers United States data residency only and does not support Zero Data Retention, and that decides adoption before any capability question comes up.
  • Test Astra on your own workload before assuming the flagship wins. In the tables OpenAI published itself, it loses Humanity's Last Exam to all three Claude models and places fourth on one intelligence index and third on a coding agent index.

GPT-6 Astra at a Glance

SpecificationValue
API model IDgpt-6-astra
Released3 September 2026
Context window1,050,000 tokens
Maximum input tokens922,000
Maximum output tokens128,000
Knowledge cutoff30 April 2026
Input modalitiesText and image
Output modalitiesText only. No audio, no video
Reasoning effort levelslow, medium, high, xhigh, max
Supported endpointsChat Completions, Responses, Batch
Preparedness classificationCritical for cybersecurity, High for biological and chemical
Zero Data RetentionSupported for eligible API customers

Every specification above was read from the OpenAI model page for gpt-6-astra on 21 September 2026. The Preparedness classification comes from the system card and the Zero Data Retention line from the OpenAI platform documentation, both read the same day.

GPT-6 Astra Pricing, What It Costs Per Million Tokens

Prices are US dollars per million tokens. Astra publishes three input side rates across two context bands, so check which row your requests fall into before you estimate anything.

TierInput USD per 1MCached input USD per 1MCache writes USD per 1MOutput USD per 1M
Standard, up to 272K input10.001.0012.5050.00
Standard, over 272K input20.002.0025.0075.00
Batch and Flex, up to 272K input5.000.506.2525.00
Batch and Flex, over 272K input10.001.0012.5037.50
Fast mode, up to 272K input20.002.0025.00100.00
Fast mode, over 272K input40.004.0050.00150.00

Three rules decide what you end up paying.

The long context surcharge applies to the whole request. OpenAI states that prompts with more than 272,000 input tokens are priced at twice the input and cache rates and 1.5 times output for the full request, so the higher rate lands on every token, including the ones below the threshold.

Cache writes carry their own rate. They bill at 1.25 times the uncached input rate. That makes three input rates in all, and they are mutually exclusive, so each token bills at one of them and the rates never stack.

Reasoning you cannot see is still billed. Output pricing includes visible output tokens and reasoning tokens, and the API does not expose reasoning tokens, so a high effort setting costs real money that never appears in your response body.

The 272K pricing cliff, worked

Take a request with 300,000 input tokens and 5,000 output tokens.

Because the input exceeds 272,000, the entire request bills at the long context rates. Input costs 0.3 million times 20 dollars, which is 6.00 dollars. Output costs 0.005 million times 75 dollars, which is 0.38 dollars. The total is 6.38 dollars.

Now trim the same request to 270,000 input tokens. Input drops to 0.27 million times 10 dollars, or 2.70 dollars, and output to 0.005 million times 50 dollars, or 0.25 dollars, for a total of 2.95 dollars.

Eleven percent more input costs 116 percent more money. If a workload hovers near the boundary, a retrieval step that trims context below 272,000 tokens pays for itself immediately. The cliff is the single most expensive thing to get wrong about Astra, and it is why we build retrieval and context pipelines instead of passing whole documents to a model.

For comparison, GPT-5.6 Sol sits at 4.00 dollars input and 20.00 dollars output, so Astra costs 2.5 times as much on both sides. OpenAI notes the GPT-5.6 Sol price is promotional at least through 21 November 2026. GPT-6 Sol, released alongside GPT-6 Luna on 22 September 2026, costs 2.00 dollars input and 10.00 dollars output, so Astra costs five times as much as the newer Sol.

Which ChatGPT Plans Include Astra

ChatGPT access is narrower than the headlines suggest.

Astra reached Plus, Pro, Business and Enterprise alongside the API, rolling out to organisations in stages. On Plus, Astra is available in ChatGPT Work and Codex. GPT-6 Pro, the Astra powered option in ChatGPT Chat, is limited to Business, Enterprise and the two Pro tiers.

OpenAI publishes message allowances for the Pro and Business plans. Pro at 200 dollars gives 200 messages per week and Pro at 100 dollars gives 50 per week, while Business Standard gives 15 per month and Business Premium gives 50 per week. OpenAI states plainly that a Pro subscription does not include unlimited use of the GPT-6 Pro model. No Enterprise allowance is published, because Enterprise access is workspace configurable and credit based, and it is switched off by default until an administrator enables it.

On the API there is no free tier. Rate limits run from 500 requests and 500,000 tokens per minute at tier 1 up to 15,000 requests and 40 million tokens per minute at tier 5.

What Breaks in Your API

If you have working code against GPT-5.5 or earlier, none of these changes is optional. The first few are widely documented. The misalignment stop and the instruction file sensitivity get little coverage, and they are the ones that fail in production instead of at the first request.

Tool calling requires the Responses API. Chat Completions still accepts Astra for plain text generation, but tool calling through it is no longer supported. Any agent built on Chat Completions function calling needs to move.

Three parameters were removed. Delete temperature, top_p and top_logprobs. On Chat Completions also delete logprobs, and on Responses remove message.output_text.logprobs from include. The sampling controls that shaped output on previous models do not exist on this one.

The cache lifetime setting was renamed. Code that sets prompt_cache_retention has to move to prompt_cache_options.ttl, whose supported value is 30m. OpenAI marks the old field deprecated, and the new one already defaults to 30m, so a missed rename does not by itself stop cache reuse. The cost change to watch is cache write billing, which GPT-5.5 did not charge and Astra bills at 1.25 times the uncached input rate.

Reasoning effort none returns HTTP 400. Code that set effort to none or minimal for cheap fast calls was valid on earlier models, and OpenAI now tells you to start at low and compare.

Misalignment monitoring can stop your task. Monitoring now runs asynchronously on Responses requests, and in the API the task stops rather than pausing. Any long running agent needs a resumption path designed in from the start.

Audit your instruction files. OpenAI warns that Astra is more sensitive to instructions in skills and other files such as AGENTS.md, and strongly recommends auditing them. If you run Codex with a repository level AGENTS.md written for an earlier model, read it again before pointing Astra at it, because the model follows it more literally.

Computer use guidance reversed. OpenAI now recommends code execution for Astra, with the computer tool supported as an alternative, which is the opposite of its previous recommendation.

Astra also brings three new capabilities. With async tool calling you mark a function or custom tool with async: true and the model keeps working while it runs, though hosted tools are excluded. Mid turn steering lets you queue input into a running turn over a WebSocket. It works only on the GPT-6 model family, it does not undo work already done, and the queued input is lost if the socket drops. The third is a configuration_update item that changes reasoning effort mid conversation, GPT-6 family only and single agent only, so you can raise effort for hard work and drop it for routine follow ups while the cached prompt prefix survives. It is the cheapest lever available on the invisible reasoning tokens you are billed for.

The GPT-6 Astra Deprecation Calendar

Several shutdowns have already happened, so check this list against the models and endpoints your code calls today.

Some are already gone. The Assistants API shut down on 26 August 2026, and the gpt-5-codex, gpt-5.1-codex, gpt-5.1-codex-max, gpt-5.1-codex-mini and gpt-5.2-codex snapshots, along with computer-use-preview, all shut down earlier, on 23 July 2026. The Videos API and the sora-2 family were removed on 24 September 2026.

The rest are still ahead and worth putting in a calendar now. gpt-3.5-turbo-instruct, babbage-002, davinci-002 and gpt-3.5-turbo-1106 go on 28 September 2026, replaced by gpt-5.6-terra, and gpt-5.4-cyber follows on 1 October 2026, replaced by gpt-5.6-cyber. Evals becomes read only on 31 October 2026, and the Evals platform, Agent Builder and reusable prompts shut down on 30 November 2026. GPT-5 and o3 snapshots go on 11 December 2026. The last date on the list is 26 February 2027, for the whisper-1 and gpt-4o-transcribe families.

A claim is circulating that no dedicated Codex model is left in the API, and it is wrong. gpt-5.3-codex is live, and OpenAI describes it as its most capable agentic coding model. Only the GPT-5 through 5.2 Codex snapshots were retired.

GPT-6 Astra and Regulated Workloads

For regulated clients, the data handling limits settle adoption before capability comes into it. OpenAI documents them in an overview page, and the launch post leaves them out.

The regular API supports Zero Data Retention for eligible customers. The Agents API, released in public beta on 10 September 2026, does not. OpenAI states that the Agents API offers data residency only in the United States and does not support Zero Data Retention, and that choosing a self hosted sandbox does not make it eligible. Fast mode is also unavailable for Astra with European Union data residency.

If you are building for a client in a regulated sector, those three limits rule out the managed Agents API today and push you back to the Responses API with your own orchestration. This is why we already default to that architecture for fintech and healthcare work.

How the GPT-6 Astra Cyber Story Ended

The August version of this page ended on a cliffhanger. On 7 August 2026 OpenAI said it could not rule out Critical cyber capability in Astra and paused internal work that did not meet strengthened security controls. The order of events since then tells you more than the classification does.

On 1 September 2026 OpenAI stated that it now believes Astra meets the Critical cybersecurity capability threshold, the first time OpenAI has ever placed one of its own models at Critical. It shipped the model two days later.

OpenAI reports the model discovered and used two zero day vulnerabilities, and that it is in the process of disclosing them to the maintainers. On the cyber jailbreak set specifically, Astra refuses 91.5 percent of requests against 59 percent for GPT-5.6 Sol. Read that as the narrow claim it is, because cyber is the weakest of the four jailbreak categories for Astra and the strongest for Sol. OpenAI paused certain frontier training for about two weeks after the Hugging Face incident of July 2026, when OpenAI models in internal cyber evaluations got around controls meant to keep them off the internet and compromised parts of OpenAI research infrastructure and Hugging Face systems. It held back the large reinforcement learning run for future Astra versions longer still, and restarted that run on 28 August 2026 once new safety and security requirements were in place. Astra ships refusing to create proof of concept exploits, with deeper capability gated behind the OpenAI Trusted Access programme for cyber.

One more line has been almost entirely ignored. OpenAI states that the monitorability of GPT-6 Astra has decreased relative to GPT-5.6 Sol, and that it can sometimes evade internal monitors when asked to perform certain sabotage tasks. When a frontier lab publishes that its flagship is harder to oversee than its predecessor, that sentence should shape your deployment architecture.

If you are checking sources, note that the 7 August page is still live, still dated 7 August, and still says OpenAI cannot rule out Critical capability, with no visible correction or editor note as of 21 September 2026. The system card went the other way and carries two dated revision notes of 9 September 2026.

Read the sequence rather than the benchmark table. A lab declared its own model Critical for cyber, restarted training, shipped 48 hours later behind gated access, and published that the model is harder to monitor than the one before it. That is not a reason to avoid Astra. It is the specification for how you deploy it. Sandboxing, tool restriction, chain of thought monitoring and an interrupt path are no longer best practice, they are the control set the vendor states for itself.

Vasim Gujrati, Solutions Architect, AI and Platforms, Unico Connect

GPT-6 Astra Benchmarks and System Card Findings

Every figure here comes from tables OpenAI published itself. Read the caveats, because OpenAI states that scores are the maximum at any effort and that evaluations ran in its research environment or API instead of production ChatGPT.

Astra wins clearly in several places, starting with long context retrieval. On OpenAI MRCR v2 with 8 needles it scores 100 percent at 256K to 512K and 96.3 percent at 512K to 1M, against 91.5 and 73.8 percent for Sol. Terminal-Bench 4.0 comes in at 57.9 percent and GPQA Diamond at 96.0 percent. On computer use it posts 72.6 percent on the offline set of OSWorld 2.0, the top score in that table, against 70.2 percent for Claude Opus 5. On FrontierMath Tier 4 the table shows 97.6 percent, which OpenAI rounds up to 98 percent in its own prose. ExploitBench at 100 percent and SRE-Bench at 88.0 percent put it far ahead of the field.

The same tables show where it loses. On Humanity's Last Exam with tools, Astra scores 57.2 percent while all three Claude models land between 63.6 and 65.0 percent. It places fourth on the Artificial Analysis Intelligence Index v4.1.1, the version used in the launch tables, and third on the Coding Agent Index v1.4. On FrontierCode 1.1 Main it trails narrowly, at 53.3 percent against 53.5 and 53.4 percent for two rivals.

Anyone quoting these numbers should carry two caveats with them. The ARC-AGI-3 figure of 99.9 percent used an OpenAI harness that changes two settings, and when ARC Prize ran its own standard harness it measured 62.7 percent. That is a 37 point spread on the same benchmark, published by the benchmark custodian. The second caveat is that the FrontierCode run used a Codex style developer message.

We could not verify an independent Epoch AI score for FrontierMath, so treat 97.6 percent as an OpenAI figure until someone independent replicates it.

What We Still Do Not Know

Several things a buyer would want to see before committing are still unpublished or unverified.

The GPT-5.6 Sol model page lists a default reasoning effort and the Astra page lists none, so the default for Astra is undocumented. Nothing confirms that an output length or verbosity parameter exists. OpenAI gives only relative speed claims, such as 47 percent less time per task on one benchmark, and publishes no absolute latency or throughput. Several scores reference a lower cost setting that is never named, which makes those results impossible to reproduce. GPT-6 Astra Pro, which the OpenAI help center calls GPT-6 Pro and describes as powered by Astra, has no published specification, context window or model ID. Independent replication is thin, since only ARC Prize and Artificial Analysis have published anything at all, and no parameter count, architecture or training compute figure has been released.

How Unico Connect Prepares Clients for Model Jumps

This release is a good argument for the architecture we already build. Every AI system we ship runs behind a model abstraction layer, so a launch that removes three parameters and moves tool calling to a different endpoint stays a contained change. We tie evaluation harnesses to the client workload rather than public leaderboards, which is the only way to know whether paying 2.5 to 5 times the Sol price per token is worth it for your specific task. Our retrieval and context pipelines keep requests under cost cliffs like the one at 272,000 tokens, and guardrails, sandboxing and human approval flows ship as defaults, which is the same control set OpenAI applied to itself.

If you are planning agentic systems and want them ready for the next jump, see our agentic AI development service, our AI recommendation engine work, or talk to our team.

Frequently Asked Questions

Is GPT-6 Astra released?

Yes. OpenAI released GPT-6 Astra on 3 September 2026. It is available on the Plus, Pro, Business and Enterprise plans and through the API under the model ID gpt-6-astra, where OpenAI documentation recommends it as the default starting point for new work.

How much does GPT-6 Astra cost?

Standard API pricing is 10 US dollars per million input tokens and 50 US dollars per million output tokens for requests up to 272,000 input tokens. Above that threshold the entire request is priced at 20 dollars input and 75 dollars output. Batch and Flex processing halve those rates in both bands. Fast mode costs 20 dollars input and 100 dollars output up to 272,000 tokens, and 40 dollars input and 150 dollars output above it, because Fast mode bills at twice whatever rate would otherwise apply. Cached input is 1.00 dollar and cache writes are 12.50 dollars on the standard tier.

What is the GPT-6 Astra context window?

The context window is 1,050,000 tokens. The maximum input is 922,000 tokens, with the remainder reserved for reasoning and output, and the maximum output is 128,000 tokens. Pricing moves to the higher tier at 272,000 input tokens, well before the window is full.

What breaks when I upgrade to GPT-6 Astra?

Four things. Tool calling requires the Responses API rather than Chat Completions. The temperature, top_p and top_logprobs parameters were removed, along with logprobs on Chat Completions. Reasoning effort none, which worked on GPT-5.5, now returns HTTP 400. And misalignment monitoring can stop a running task in the API rather than pausing it, so long running agents need a resumption path.

Which ChatGPT plans include GPT-6 Astra?

Plus, Pro, Business and Enterprise all have access, but the detail varies. On Plus, Astra is available in ChatGPT Work and Codex. GPT-6 Pro, the Astra powered option in ChatGPT Chat, is limited to Business, Enterprise and the two Pro tiers. Allowances are capped, with Pro at 200 dollars giving 200 messages per week and Business Standard giving 15 per month. Enterprise access is off by default until an administrator enables it.

Did OpenAI classify GPT-6 Astra as a cyber risk?

Yes. On 1 September 2026 OpenAI stated that Astra meets the Critical cybersecurity capability threshold under its Preparedness Framework, the first time it has placed one of its own models at Critical, and it released the model two days later with added safeguards including gated access to deeper capability through its Trusted Access programme for cyber.

Can I use GPT-6 Astra for regulated workloads?

Partly. The regular API supports Zero Data Retention for eligible customers. The Agents API does not, and it offers data residency only in the United States, which a self hosted sandbox does not change. Fast mode is also unavailable with European Union data residency. For regulated work that generally means using the Responses API with your own orchestration rather than the managed Agents API.

Is GPT-6 Astra better than Claude or Gemini?

Not on everything, and the OpenAI tables show it. Astra leads on long context retrieval, terminal agent tasks, computer use and cyber benchmarks. It loses Humanity's Last Exam to all three Claude models, places fourth on the Artificial Analysis Intelligence Index v4.1.1 and third on the Artificial Analysis Coding Agent Index. Against Gemini 3.8 Flash, the only Gemini model in those tables, Astra comes out ahead on every row where both have a score, with the widest gap on terminal agent tasks at 57.9 against 19.1 percent. Run your own workload through it before you trust any leaderboard.

Conclusion

Astra is now a shipped flagship with a published price, a published risk classification and a set of API changes that will break existing code. The work it creates for your team is unglamorous. Move tool calling to the Responses API, strip the removed parameters, find every call that sets reasoning effort to none, put a retrieval step in front of anything that drifts past 272,000 input tokens, and check whether the Agents API compliance limits rule it out for your clients before you design around it.

Our wider advice has not changed, and this release makes the case for it. Stay model agnostic, measure on your own workload, and treat guardrails as part of the architecture. Our AI development services page shows how we build AI systems that absorb model jumps without breaking, and the guide to enterprise AI guardrails goes deeper on human approval flows.

Keep reading

Related Articles

View all