Anthropic released Claude Opus 5.5 on 22 September 2026. List prices for input and output tokens are 20% below Opus 5, while cache-read prices are 60% lower. Anthropic reports a 40% reduction for typical workloads because it combines those lower prices with a claim that the new model uses fewer tokens per task.
That headline does not justify an automatic production switch. If you already use Opus 5 for long-running agents, coding, document processing or knowledge work, run a controlled comparison on real, anonymised Hungarian-language tasks or equivalent local-language workflows elsewhere in CEE. Keep prompts and cache settings constant, test comparable effort levels, then compare total cost per successful task. High-volume systems using smaller model tiers have a good reason to wait for Sonnet 5.5 and Haiku 5.5 data.
How much cheaper is Opus 5.5?
According to Anthropic's 22 September announcement, Opus 5.5 costs $4 per million input tokens and $20 per million output tokens. Cache reads cost $0.20 per million tokens and cache writes cost $5.
| Charge | Opus 5 | Opus 5.5 | Change |
|---|---|---|---|
| Cache reads | $0.50 | $0.20 | 60% lower |
| Input tokens | $5 | $4 | 20% lower |
| Output tokens | $25 | $20 | 20% lower |
| Cache writes | $6.25 | $5 | 20% lower |
The table uses US dollars per million tokens. Source: Anthropic, 22 September 2026.
Anthropic says Fast mode, offered in Claude Code and on the Claude Platform, provides up to 2.5 times the speed at $8 per million input tokens and $40 per million output tokens. That is twice the standard Opus 5.5 rate. It is relevant when response time has measurable business value, but it should be tested as a separate pilot configuration.
Why won't every company save 40%?
The difference between 20% and 40% comes from two effects. Most list-price components fell by 20%. Anthropic also says Opus 5.5 consumes fewer tokens to complete a task. The second effect needs to be measured on the buyer's own workload, particularly for Hungarian text, because the company does not provide Hungarian-language quality or tokenisation data.
An illustrative calculation assumes one agent run uses 2 million cache-read tokens, 0.1 million input tokens and 0.02 million output tokens, with no cache write. The calculation also assumes identical token use on both models, so Opus 5.5 receives no additional efficiency benefit.
* Opus 5: $1 for cache reads, $0.50 for input and $0.50 for output, totalling $2. * Opus 5.5: $0.40 for cache reads, $0.40 for input and $0.40 for output, totalling $1.20.
In this cache-heavy example, price changes alone reduce the cost by 40%. A drafting workload dominated by new input and long output would see roughly a 20% reduction until lower token use appears in the buyer's own test.
The useful purchasing metric is therefore cost per successful task. The guide to proprietary enterprise models compares additional pricing, cache and cloud factors.
Which workflows justify a pilot now?
The strongest candidate is a process that already runs on Opus 5, repeatedly reuses a large context and has a clear definition of success. Examples include a long development session, an agent that reviews several documents, or a company knowledge base that receives regular updates.
| Buyer situation | What to do now | Why |
|---|---|---|
| Opus 5 agent with heavy cache use | Run a controlled pilot | Cache-read pricing fell by 60% |
| Assistant producing long answers | Measure output and human correction time separately | Output pricing fell by 20%; further savings remain workload-dependent |
| Workflow already performing well on Fable 5.1 | Test cost on the same tasks and quality threshold | Anthropic claims Fable 5.1-level performance on most work. In its own HAProxy test, Opus 5.5 cost 51% less while both models passed nearly all regression tests. This is a vendor test, so confirm quality on your own tasks |
| Team using Microsoft 365 Copilot | Check the licence, region and whether the model has reached the tenant | Microsoft says availability varies by licence, access and region |
| High-volume system using a smaller model | Wait for Sonnet 5.5 and Haiku 5.5 data | Anthropic says both will follow in the coming weeks |
For a Hungarian workflow, include inflected terms, declined names, long compound sentences, tables and local business documents. These may include supplier contracts, public-procurement attachments or NAV Online Számla exports when the process handles them. For another CEE market, substitute the relevant local tax and procurement documents. A practical set is 30–50 real, anonymised tasks. Track task success, human correction time, tool calls, elapsed time and token use by category.
For production agents, the model is only one part of the system. Access control, tool permissions, logging and approval design are covered on the AI agent development page.
Which purchasing channel fits your data requirements?
Anthropic's Opus product page lists Opus 5.5 on the Claude Platform, AWS, Google Cloud and Microsoft Foundry. The choice should follow the required data location, existing cloud contracts, access controls and operations tooling.
The AWS announcement says Amazon Bedrock provides zero data retention support by default and keeps data within AWS infrastructure with regional data residency. Buyers still need to check the specific AWS region and contractual terms applying to their service configuration.
Microsoft began rolling out Opus 5.5 and OpenAI's GPT-6 Sol on the same day across Word, Excel, PowerPoint, Chat, Cowork and Copilot Studio. Its rollout notice says availability may vary by licence, access and region. A Hungarian tenant should therefore be checked against the Microsoft AI at Work Roadmap and Microsoft Copilot release notes before a project plan assumes access.
Anthropic's product page states that US-only inference is available at 1.1 times the standard input and output prices, but it does not state an EU-only inference option. Where data location is a contractual requirement, select the channel only after reviewing its data-processing terms. The broader architecture decision is covered in the guide to on-premise and cloud AI.
If you are comparing several models and cloud channels, enterprise AI platform design helps separate model, data and operations decisions. Enterprise AI platforms
How much weight should you give the benchmarks?
Anthropic reports leading Opus 5.5 results across several coding, computer-use and knowledge-work benchmarks. The company also says benchmark margins have become a less reliable guide to real-world differences at this capability level.
The test conditions are not uniform. Several Opus 5.5 results use maximum or xhigh effort, while some competitor figures are those reported by OpenAI. When built-in safeguards intervened, cybersecurity, biology and frontier LLM development tasks were completed by older Claude models. Anthropic says this likely reduced the Opus 5.5 scores. In Zapier's AutomationBench runs, safeguard interventions counted as failures.
The FrontierCode description gives Opus 5.5 a score of 54.6% at medium effort, while the table reports 54.4% at maximum effort. That 0.2-point difference is too small to support an effort-setting decision by itself. A pilot should test more than one effort level and evaluate task quality together with cost per successful task.
How should you compare Opus 5.5 with Opus 5?
- Select 30–50 anonymised tasks from the production workflow, including the Hungarian-language cases that currently require correction.
- Run identical prompts on Opus 5 and Opus 5.5 with the same cache settings and several comparable effort levels.
- Use the explicit
claude-opus-5-5model identifier and record the tested configuration. - Track input, output, cache reads, cache writes, tool calls, elapsed time and human correction time separately.
- Compare total cost per successful task. Price per token does not capture failed attempts or review work.
- Before production deployment, confirm the chosen channel's region, retention terms, licence and access conditions.
Who benefits from switching to Opus 5.5 now?
A pilot is most relevant to companies already using Opus 5 for long agent runs with heavy cache use. In the illustrative example, the new list prices reduce the cost of a run by 40% without assuming lower token use.
If you use Fable 5.1, test Opus 5.5 for cost against the same quality threshold. If a smaller model already meets the required quality, test Opus 5.5 only against a specific shortcoming. Teams without an existing Opus-class workload can wait for Sonnet 5.5 and Haiku 5.5 data. A production switch is justified when Opus 5.5 delivers the required quality at a lower total cost on the company's own tasks.
Chart data as a table
| Cost of an illustrative agent run | $ |
|---|---|
| Opus 5 | 2 |
| Opus 5.5 | 1.2 |
