OpenAI released GPT-6 Astra on September 3, two days after Anthropic released Claude Fable 5.1 on September 1. The list price is identical: $10 per million input tokens and $50 per million output tokens on both models. Cache reads cost $0.25 on Fable 5.1 and $1 on Astra, so Fable 5.1's cache reads are a quarter of the price. On benchmarks there is no agreement: in OpenAI's own table Astra leads Terminal-Bench 4.0 (57.9% against 55.8%), while on the Artificial Analysis Intelligence Index Fable 5.1 is ahead (66 points to Astra's 61). Neither model leads everywhere; the decision comes down to your workload profile and your own eval set.
What is new in Astra compared to GPT-5.6 Sol?
From OpenAI's announcement this article takes five areas: computer use, professional work, agentic coding, mathematics and science, and staying within scope. Most numbers come with the value of the predecessor GPT-5.6 Sol; a Fable 5.1 value appears where OpenAI's table carries one.
In computer use, OpenAI also simulated time per task on OSWorld 2.0. Astra scores 72.6% at roughly 40 minutes per task, Sol 65.7% at roughly 75 minutes; per OpenAI that is about 47% less time per task. The updated Codex harness together with Astra completes Mind2Web tasks 1.9x faster than the current Sol-based Codex. These are two separate measurements: the first measures the model alone, the second the model and the harness together. For Fable 5.1, OpenAI publishes no value in the OSWorld row, and its footnote says the Claude models were run on the official settings. Anthropic's 77.9% partial Fable 5.1 result, measured on its own settings, was therefore produced under different settings than Astra's 72.6%; the two numbers cannot be compared.
In professional work, on AutomationBench, which measures business workflows, Astra scores 41.4%, Sol 18.1% and Fable 5.1 31.4%. On BenchCAD, Astra scores 95.9% and Fable 5.1 84.3%; per OpenAI, the Claude scores here reflect the three modifications described in the Fable 5.1 System Card.
In agentic coding, on Terminal-Bench 4.0, Astra scores 57.9%, Sol 37.3% and Fable 5.1 55.8%. On DeepSWE v1.1, Astra scores 74.1% and Fable 5.1 67.4%; on FrontierCode 1.1 Extended, 64.5% and 63.6% respectively. In Codex, Astra keeps notes across context windows as an experimental feature, and it asks clarifying questions asynchronously while continuing the work that does not depend on the answer.
In mathematics and science, the FrontierMath Tier 4 (v2) row shows 97.6% for Astra and 87.8% for Fable 5.1 (OpenAI's prose rounds to 98%). On Terminal-Bench Science 0.1, Astra scores 64.6% and Fable 5.1 52.6%. Astra helped improve the bound on gaps between primes from 240 to 186. On the tool-assisted Humanity's Last Exam the order is reversed: 65.0% for Fable 5.1 and 57.2% for Astra.
In the test that measures staying within scope, Sol without production safeguards went beyond the authorized target in 48% of cases, Astra in 0%.
Which model leads the benchmarks, and in whose table?
The two developers give two numbers for the same row. Anthropic's announcement puts Sol at 19.6% on AutomationBench, OpenAI's table at 18.1%; Fable 5 is 17.1% per Anthropic and 17.4% per OpenAI. On Fable 5.1's Terminal-Bench 4.0 value (55.8%) the two companies agree.
Chart data as a table
| Benchmark | GPT-6 Astra | Claude Fable 5.1 | GPT-5.6 Sol |
|---|---|---|---|
| AutomationBench | 41.4% | 31.4% | 18.1% |
| Terminal-Bench 4.0 | 57.9% | 55.8% | 37.3% |
| Terminal-Bench Science 0.1 | 64.6% | 52.6% | 22.4% |
| DeepSWE v1.1 | 74.1% | 67.4% | 72.7% |
| BenchCAD | 95.9% | 84.3% | 83.3% |
| FrontierMath Tier 4 (v2) | 97.6% | 87.8% | 83.0% |
| GPQA Diamond | 96.0% | 93.7% | 94.6% |
| Humanity's Last Exam (w/ tools) | 57.2% | 65.0% | Not published |
The chart's Sol values from OpenAI's table: DeepSWE 72.7%, BenchCAD 83.3%, Terminal-Bench Science 22.4%, FrontierMath Tier 4 83.0%; on GPQA Diamond Astra scores 96.0%, Fable 5.1 93.7% and Sol 94.6%.
The Artificial Analysis comparison is the third view: on the Intelligence Index, Fable 5.1 (max effort) scores 66 points and Astra (max) 61 points, and it also rates Fable 5.1 as cheaper: its blended price is $7.17 per million tokens against Astra's $7.70 per million tokens, calculated at a 7:2:1 cache-read/input/output ratio. OpenAI's own table carries the same index in version v4.1.1 at 61.2 and 65.7 points, so this order appears in OpenAI's table too. Read the footnotes of the benchmark table as well, and measure on your own tasks.
How much do the two models cost on the API?
The list price is identical on both models: $10 per million input tokens and $50 per million output tokens, with batch processing at $5 and $25. The first difference is cache reads: $1 per Astra's model page, $0.25 per Fable 5.1's documentation, that is, a tenth of the base price on Astra and a fortieth on Fable 5.1. Cache writes cost $12.50 on Astra, and on Fable 5.1 $12.50 for a 5-minute cache and $20 for a 1-hour cache. A long-running agentic session re-reads its cached prefix at every step, so the cache-read price carries the difference.
Chart data as a table
| USD per 1M tokens | GPT-6 Astra | GPT-5.6 Sol | Claude Fable 5.1 |
|---|---|---|---|
| Input | $10 | $4 | $10 |
| Cache read | $1 | $0.40 | $0.25 |
| Output | $50 | $20 | $50 |
Sol's cache reads cost $0.40 per million tokens.
The second difference is long context. Astra's model page prices prompts above 272K input tokens at 2x the input and cache rates and 1.5x the output rate for the full request; the long-context column of OpenAI's pricing page reads $20, $2, $25 and $75. Fable 5.1 bills at standard pricing across the full 1M-token context window. Fast mode on Astra promises up to 2x the speed at 2x the price ($20 and $100), but it is unavailable with EU data residency. Regional processing at OpenAI carries a 10% uplift for models released on or after March 5, 2026.
OpenAI also publishes cost-per-task estimates: on Terminal-Bench 4.0 Astra is about 63% cheaper than Fable 5.1, on Terminal-Bench Science about 31% cheaper, and on BenchCAD about 86% cheaper. These are OpenAI's estimates; the Artificial Analysis blended price points the other way. One number measures token price, the other token consumption. GPT-5.6 Sol is listed at $4 input and $20 output, at a promotional price through at least 21 November 2026; Astra costs two and a half times that.
For a live system, measuring the two models side by side, testing the switch, and rolling it out is an operations task — we carry it on a monthly plan too. Managed AI Operations
Who gets access, and how does each model handle data?
Per the announcement, Astra opened on day one to a limited set of organizations and arrives over the coming days for ChatGPT Plus, Pro, Business and Enterprise plans, the OpenAI API, Microsoft Azure and AWS Bedrock. The Help Center article is more specific: in Chat, Astra appears as GPT-6 Pro on the Pro $100, Pro $200, Business and Enterprise plans, and the Plus plan does not include it in Chat. The usage allowance: 200 messages per week on Pro $200, 50 per week on Pro $100, 15 per month on Business Standard and 50 per week on Business Premium, the last three from an allowance shared with GPT-5.6 Sol Pro. In Enterprise workspaces access is off by default and an administrator enables it.
ZDR is inverted between the two models: Astra grants zero data retention to eligible API customers, Fable 5.1 ties it to express authorization. Fable 5.1 comes with 30-day data retention, its platforms are the Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry, and its text output carries Anthropic's statistical watermark on every platform.
The operational difference is in oversight. Based on the System Card, OpenAI placed Astra at the Critical cybersecurity level under the Preparedness Framework, the company's own safety framework. Per the safety overview, it is the first OpenAI model at that level, and misalignment monitoring runs on all tool-using inference in Astra's external deployment. OpenAI's announcement adds that if that check stops a task, in ChatGPT and Codex the user can review the step; in the API the task stops. Your calling code has to handle that stop: if a long-running agentic session halts mid-way through its steps, your system has to handle the restart, saving partial results, and escalation to a human. Astra refuses more advanced cybersecurity requests such as writing proof-of-concept exploits; OpenAI promises less restrictive safeguards through the Daybreak program in the coming weeks. On Gray Swan's 1810 curated prompt-injection attacks, the safeguards-enabled Astra showed an attack success rate of 8.5%, Sol 27.0%. The same System Card describes that Astra's chain of thought became shorter and less informative, so it is harder to monitor than Sol.
What should you test before switching?
Astra's model id is gpt-6-astra. Chat Completions, Responses and Batch are supported endpoints; Realtime, Assistants, fine-tuning and embeddings are not supported. reasoning.effort takes low, medium, high, xhigh or max. The context window is 1,050,000 tokens, with a maximum input of 922,000 and a maximum output of 128,000 tokens; the training data runs through April 30, 2026. Astra's capabilities and the full table measured against GPT-5.6 Sol are covered in a separate article.
On Fable 5.1, migration touches three API changes: forced tool use returns a 400 error, earlier models cannot read Fable 5.1's thinking blocks, and editing earlier turns invalidates thinking blocks. The three compatibility points and the cache math are detailed in our article on Fable 5.1.
Before switching, measure five things on your own eval set:
- Token consumption per task. OpenAI's per-task cost advantage is its own estimate; at an identical list price it only holds if the model uses fewer tokens on your tasks too.
- Cache hit ratio. If most of your input is cache reads, the gap between Fable 5.1's $0.25 and Astra's $1 decides.
- Share of long requests. Work out what percentage of your requests exceed 272K input tokens; Astra bills those requests entirely at the higher price.
- Stops and refusals. On Astra the safety check can stop a task in the API; measure how often and where.
- Compatibility. From Fable 5, the three API changes; from Sol, the pricing rule and the Fast mode EU limitation.
Which one should you choose?
Astra is indicated if your system is built on computer use and time per task matters, if you run a terminal-based coding agent, or if you have to offer ZDR to your customer within the OpenAI ecosystem. Fable 5.1 is indicated if the workload is heavily agentic and re-reads a large cached prefix at every step, or if you want to use the full 1M-token context window at a uniform price.
There is no forced move on either side. GPT-5.6 Sol stays on the price list, with its promotional price valid through at least 21 November 2026. For most workloads, Anthropic's documentation recommends Claude Opus 5 as the starting point. Narrow down by your workload profile first, then measure both candidates on your own eval set with identical prompts and an identical cost budget; that is where the decision is made.

