OpenAI released GPT-6 Astra on September 3, the successor to July's GPT-5.6 Sol. According to the company's announcement, the model posts the best published results in computer use, browsing, software engineering, cybersecurity, science, and professional work. OpenAI leads with computer use: in its OSWorld 2.0 latency simulation, Astra scores 72.6% at roughly 40 minutes per task, against Sol's 65.7% at roughly 75 minutes. Token prices are $10 per million input tokens and $50 per million output tokens, compared with $4 and $20 for Sol. Astra is also the first OpenAI model to reach the Critical cybersecurity level under the company's Preparedness Framework, its safety framework. We walk through what is new, what it costs, and who gets it.
What is new in computer use?
OpenAI gives computer use — operating a browser and desktop applications — its own section and calls Astra "the world's best computer use model". In the OSWorld 2.0 latency simulation, Astra reaches a higher score in about 47% less time per task than Sol. This is OpenAI's own simulation; no independent measurement of the speed figure exists on release day.
OpenAI measures the Codex harness update separately: the model and the new harness together complete tasks 1.9x faster on the Mind2Web benchmark than the current Sol-based Codex. The two figures are separate measurements: one covers the model alone, the other the model plus the harness.
What does it bring to professional work and coding?
For a business reader the key row is AutomationBench, which measures business workflows: Astra scores 41.4%, Sol 18.1%; OpenAI's table puts Claude Fable 5.1 at 31.4% and Opus 5 at 26.9%.
Chart data as a table
| Benchmark | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|
| AutomationBench | 41.4% | 18.1% |
| Terminal-Bench 4.0 | 57.9% | 37.3% |
| Terminal-Bench Science 0.1 | 64.6% | 22.4% |
| OSWorld 2.0 (offline, partial) | 72.6% | 65.7% |
| ScreenSpot-Pro | 92.7% | 76.9% |
| FrontierMath Tier 4 (v2) | 97.6% | 83.0% |
| ARC-AGI-3 | 99.9% | 7.8% |
| ExploitBench | 100.0% | 78.5% |
The chart's remaining rows from OpenAI's table: OSWorld 2.0 offline set with partial scoring 72.6% (Sol 65.7%), ScreenSpot-Pro 92.7% (Sol 76.9%), GPQA Diamond 96.0% (Sol 94.6%), ARC-AGI-3 99.9% (Sol 7.8%), ExploitBench 100.0% (Sol 78.5%).
In coding, Terminal-Bench 4.0 shows 57.9% for Astra, 37.3% for Sol, and, per OpenAI's table, 55.8% for Fable 5.1. By OpenAI's own estimate, Astra's API cost per task on this test is about 63% lower than Fable 5.1's (about 9% lower than Sol's).
In Codex, Astra carries notes across context windows (an experimental feature) and asks its clarifying questions asynchronously, without interrupting the work; on decisions with serious consequences it waits for your answer.
What does it bring to mathematics, science, and long context?
On FrontierMath Tier 4 (v2), Astra scores 97.6% — OpenAI rounds this to 98% in its prose — against Sol's 83.0% and Fable 5.1's 87.8%.
On Terminal-Bench Science 0.1, which measures scientific work, Astra scores 64.6%, Sol 22.4%, and Fable 5.1 52.6%. One row runs the other way: on the tool-assisted variant of Humanity's Last Exam, Astra scores 57.2% and, per OpenAI's table, Fable 5.1 scores 65.0%. In long context, OpenAI's MRCR v2 (8-needle, 512K–1M) shows 96.3% against Sol's 73.8%.
The publishers' numbers differ. On AutomationBench, OpenAI reports 18.1% for Sol, while Anthropic's announcement gives 19.6%. On OSWorld 2.0, OpenAI's table carries no value for Fable 5.1, and OpenAI notes that it scored Claude with the official settings; Anthropic, with its own settings, measured 77.9% on the partial score. The two publishers run the same benchmark with different settings, so compare numbers within one publisher.
How well does it respect scope, and how honest is it?
OpenAI built a test modeled on the Hugging Face incident to measure whether the model goes beyond the authorized target. Without production safeguards, Sol went beyond it 48% of the time; Astra did so in 0% of cases. In another evaluation, Astra is three times less likely than Sol to misrepresent its own capabilities.
Per the safety overview, Astra is also more robust than Sol to jailbreaks and prompt injections. Per the System Card, on 1,810 curated attacks from Gray Swan's IPI Arena, the attack succeeded against the safeguards-enabled Astra 8.5% of the time, against Sol 27.0%.
OpenAI also describes two regressions. Astra's chain of thought is shorter and less informative than Sol's, so the model's monitorability has decreased, and in adversarial tests, where OpenAI pushes it to evade, it sometimes evades the internal monitors on sabotage tasks. According to Apollo Research, Astra verbalizes that it is being tested in 41.1% of samples (xhigh effort) and 50.6% (max effort), so a low misbehavior rate is weak evidence of alignment.
In the biological and chemical category, OpenAI treats the model as High as a precaution. Because of the Critical cybersecurity designation, the current release refuses advanced cybersecurity tasks such as writing proof-of-concept exploits; OpenAI plans to make less restrictive safeguards available through the Daybreak program in the coming weeks. OpenAI also monitors suspicious steps in production traffic: in ChatGPT and Codex, a paused task asks the user to review the action; in the API, the task stops.
If you are taking the model switch through a live enterprise system, our enterprise AI platforms service covers model selection, the cost model, and data handling together. Enterprise AI Platforms
What does it cost, and how does it differ from the GPT-5.6 family's pricing?
Per the model page and the pricing page, gpt-6-astra's Standard price per million tokens is $10 for input, $1 for cache reads, $12.50 for cache writes, and $50 for output. Above 272K input tokens the whole request costs more: input and cache rates double and output rises by half, to $20, $2, $25, and $75. Batch and Flex processing run at half the Standard price: $5 input and $25 output. Fast mode delivers up to twice the speed at twice the Standard price ($20 and $100); with EU data residency, Fast mode is unavailable for Astra, Standard processing applies, and regional processing carries a 10% uplift. Eligible API customers can request Zero Data Retention (ZDR).
Chart data as a table
| USD per 1M tokens | GPT-6 Astra | GPT-5.6 Sol | Claude Fable 5.1 |
|---|---|---|---|
| Input | $10 | $4 | $10 |
| Cache read | $1 | $0.40 | $0.25 |
| Output | $50 | $20 | $50 |
Sol's cache reads cost $0.40 and Fable 5.1's $0.25 per million tokens.
In the same table, Sol costs $4 for input and $20 for output. That is a promotional price: per OpenAI's July announcement, Sol launched at $5 and $30, dropped by more than 20% on August 21 and the promotional price runs at least through November 21, 2026. The GPT-5.6 family arrived on July 9 in three tiers: Sol is the flagship, Terra the balanced model, Luna the cheapest; the number marks the generation, the name the capability tier. Terra costs $2 and $12 per million tokens, Luna $0.20 and $1.20. Astra is therefore two and a half times Sol's promotional price on both lines.
Per Anthropic, Claude Fable 5.1 is also $10 and $50, but cache reads cost $0.25 against Astra's $1. The batch price ($5 and $25, the same for both models), the 30-day data retention, and ZDR by express authorization only are covered in our Fable 5.1 article.
The context window is 1,050,000 tokens, of which at most 922,000 can be input; output is capped at 128,000 tokens. The knowledge cutoff is April 30, 2026, meaning the training data runs to that date. Reasoning effort has five levels: low, medium, high, xhigh, and max.
Who gets it, and when?
The rollout started on September 3 with a limited set of organizations. Over the coming days the model becomes available through the API, Microsoft Azure, and AWS Bedrock, and in the Plus, Pro, Business, and Enterprise plans. In an Enterprise workspace the administrator enables it; at launch it is off by default.
In the ChatGPT chat interface the model appears as GPT-6 Pro on the $100 and $200 Pro plans, Business, and Enterprise. The announcement lists Plus subscribers too, but per the Help Center the Plus plan does not include GPT-6 Pro in Chat. The message limits: 200 messages per week on the $200 Pro plan and 50 per week on the $100 plan; Business Standard gets 15 messages per month and Business Premium 50 per week. The $100 Pro plan and both Business plans share their allowance with GPT-5.6 Sol Pro; on the $200 Pro plan, Sol Pro has a separate daily allowance.
OpenAI has published no separate EEA availability statement for Astra so far; the Help Center states for GPT-5.6 that eligible users in the EEA, Switzerland, the United Kingdom, and the UAE can use it. If you run several models on one shared platform, ZDR, regional processing, and switching between models can be handled in one place; we describe this on our enterprise AI platforms page.
Who should switch, and what should you test first?
Switching makes sense if you build on computer use, where the published speed gap to Sol is about 47%, or if you automate business workflows: Astra's 41.4% on AutomationBench is more than double Sol's 18.1%. Among the rows covered here, the Sol-to-Astra gap is largest on Terminal-Bench Science (22.4% → 64.6%), AutomationBench, and MRCR v2 (73.8% → 96.3%). Migration is, in the base case, a model-ID change (gpt-5.6-sol → gpt-6-astra). Waiting makes sense if Sol solves your tasks at two and a half times less cost. Sol stays on the price list with its promotional price through at least November 21, 2026, so there is no forced move. The bill and the behavior shift on three points:
- The 272K threshold. Above it the whole request gets more expensive: input rises to $20 and output to $75. Measure how many of your requests cross the line.
- The cache ratio. Cache writes cost $12.50 (1.25 times the input price) and cache reads $1; on Fable 5.1 cache reads cost $0.25. A long-running agentic session re-reads its cached prefix at every step, so this line decides which model is more expensive for your workload.
- Stops in the API. If misalignment monitoring halts a task, the API call stops. The workflow needs a retry or a human step for that case.
Measure on your own eval set; the benchmark table alone does not decide. From the Claude side the picture differs: Artificial Analysis's comparison puts Claude Fable 5.1 at 66 points and Astra at 61 on its Intelligence Index, and at a 7:2:1 cache/input/output ratio it finds Fable 5.1 cheaper ($7.17 against $7.70 per million tokens). OpenAI's own table shows 61.2 and 65.7 on the same index (v4.1.1). Our row-by-row comparison of the two models covers the rest.

