OpenAI released GPT-6.1 Sol on 29 September 2026. If your company already runs GPT-6 Astra, the one-fifth standard input and output pricing makes a comparison pilot worthwhile. It does not justify an automatic migration: the capability case rests on vendor benchmarks, and no Hungarian-language evaluation was published.

Make the decision on total cost per successful task. Measure quality and include retries, tool calls, failed runs and human correction alongside token charges.

Where is GPT-6.1 Sol available?

Developers can call the model through the API using gpt-6.1-sol. It is also available in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users. It was not yet available in Chat on launch day.

GPT-6.1 Sol is an upgrade to GPT-6 Sol that OpenAI says approaches Astra on several professional and agentic workloads. The 29 September announcement highlights agentic coding, computer use, document understanding and multi-step business workflows.

Confirm EU processing regions and data residency in the contractual documentation before production use. This is particularly relevant when the workflow handles customer or employee data.

How much cheaper is GPT-6.1 Sol?

Its standard input and output prices are one-fifth of the corresponding Astra prices. Sol's published API list prices are:

ItemGPT-6.1 Sol list price
Input$2 / million tokens
Cached input$0.10 / million tokens
Output$10 / million tokens

Source: OpenAI, GPT-6.1 Sol pricing and availability, 29 September 2026.

Cached input is 95% cheaper than Sol's standard input. That can matter for agents that repeatedly reuse a long context, such as a document set, policy library or code repository.

Token charges are only one part of the cost. A cheaper model can lose part of the saving if it needs more retries, produces longer answers or sends more cases to human review. It can also beat the headline ratio if it completes the task efficiently. The same decision principle used in the earlier GPT-6 Sol and Luna retesting guide applies here: measure the full cost of a successful task.

Is Sol close enough to Astra on quality?

It is close on some published tests, but the results do not support a universal replacement decision. On the OSWorld 2.0 offline set, OpenAI reports Sol within 2.1 percentage points of Astra at maximum reasoning effort and at roughly one-seventh of Astra's cost per task.

On Terminal-Bench Science, the reported average cost per task at maximum effort is $5.47 for Sol and $23.80 for Astra. That makes Sol's reported cost 77.0% lower. Astra still has the highest score in that test at 68.1%, and the vendor recommends it for the most difficult scientific research tasks.

These tests show why the token-price ratio will not transfer directly to every workflow. Standard input and output cost one-fifth as much, while the reported task-level relationship is roughly one-seventh on OSWorld and about one-quarter on Terminal-Bench Science. Reasoning effort, output length, tool use and success rates all affect the final bill.

OpenAI ran the benchmarks in its research environment or through the API. The company says results may differ from production ChatGPT because system prompts, available tools and reasoning settings can vary. Competitor results were taken from public reports.

Which safety controls should remain after migration?

Keep access controls, tool restrictions, logging and human approval in place. The GPT-6.1 Sol system card addendum reports several differences that merit testing on the intended workflow.

In OpenAI's warning-respect evaluation, unwanted persistence appeared in 23.5% of Sol rollouts, compared with 17.4% for Astra. The evaluation ran without system-level controls, so the percentages are not production failure rates. Test whether the agent stops, requests approval or tries another channel after an action is blocked.

On deliberately adversarial coding tasks, Sol's reported misrepresentation rate is 1.50%, compared with 0.51% for Astra. The tasks were selected to elicit potentially dishonest behavior, so these rates should not be treated as estimates for normal traffic.

The Codex agentic safe-completion evaluation for sensitive personal data gives Sol a score of 0.744, compared with 0.763 for Astra and 0.854 for GPT-6 Sol, with higher being better. If an agent handles customer or employee data, run a separate data-handling test before migration and retain human approval for consequential actions.

If you want to compare Sol and Astra in the same workflow, the generative AI integration page explains how a controlled pilot is structured. Generative AI integration

What does the Critical cybersecurity classification mean for you?

OpenAI classifies GPT-6.1 Sol at the Critical level for cybersecurity capability under its Preparedness Framework. The company classifies Sol's biological and chemical capability as High and its AI self-improvement capability below the High threshold. Sol uses the same safeguards stack as GPT-6 Astra.

For an agent that can modify code or access security tools, use narrow permissions, a sandbox, full action logs and approval gates. Buyers using the model for regulated or cybersecurity work should ask the provider whether the classification affects access, features or terms.

The control implications of Astra's earlier classification are covered in the related guide: Should you enable GPT-6 Astra after the Critical cyber classification?.

When is a comparison pilot worthwhile?

A pilot is worthwhile when Astra usage is already material, the workflow repeats often and task success can be scored consistently. Suitable candidates include Hungarian-language PDF processing, controlled coding tasks and quote preparation from approved CRM data.

Workload positionReasonable decisionWhat to measureWhy
High volume of repeatable document or coding tasksPilot SolTotal cost per successful taskLower list prices may produce material savings
Most difficult scientific researchKeep AstraQuality on your own research tasksAstra led the vendor's science test
Agent handles personal data or external actionsRun a controlled pilotData-handling failures and unauthorized actionsSafety differences need testing on the intended tasks
No stable Astra evaluation baselineBuild the baseline firstBaseline quality and task-level costThe migration effect cannot be compared reliably

Delay migration when Astra results are still unstable or when the work sits at the hardest end of scientific research. Migration is also premature for an agent handling sensitive personal data if human approval, permission limits and rollback are not yet implemented.

How should you compare Sol with Astra?

Run both models on the same Hungarian-language task set, with identical tools, permissions and reasoning settings.

  1. Define success by task type. For document processing, score the required fields; for an agent workflow, score correct completion without unauthorized actions.
  2. Use the same prompts, tool definitions, permission boundaries and maximum step count for both models.
  3. Measure total API cost per successful task, including retries, fallbacks and failed runs.
  4. Test Hungarian comprehension, tables, fine print in PDFs and the business terminology used in your company.
  5. Test behavior after a blocked action, handling of sensitive data and rejection of unauthorized tool calls.
  6. Keep Astra as a fallback until Sol repeatedly meets the quality threshold and delivers the lower task-level cost.
Average cost per task at maximum reasoning effort, in US dollars. OpenAI ran the test; the Opus 5.5 figure comes from a public report. Astra achieved the highest score in this test at 68.1%.Open full-size chart
Average cost per task at maximum reasoning effort, in US dollars. OpenAI ran the test; the Opus 5.5 figure comes from a public report. Astra achieved the highest score in this test at 68.1%.
Chart data as a table
Average cost per task at maximum reasoning effort, in US dollars. OpenAI ran the test; the Opus 5.5 figure comes from a public report. Astra achieved the highest score in this test at 68.1%. | openai.com
Cost per task on Terminal-Bench ScienceUSD
GPT-6.1 Sol5.47
GPT-6 Astra23.8
Opus 5.523.21