Claude Opus 5.5 cuts agent costs: what the new pricing changes
Anthropic says Claude Opus 5.5 costs about 40% less than Opus 5 on typical token-billed workloads. Here is what changed in token pricing, cache reads, and the economics of long-running agents.
Claude Opus 5.5 changes the economics of Anthropic’s top-tier model in two different ways. The easy-to-see change is a 20% cut to list input and output token prices. The more important claim for agent builders is Anthropic’s estimate that typical token-billed workloads cost about 40% less than Opus 5 because the new model also uses fewer tokens per task and makes cache reads much cheaper.
What actually got cheaper
Opus 5.5 vs Opus 5 list pricing
| Cost item | Opus 5.5 | Opus 5 | Change |
|---|---|---|---|
| Input / 1M tokens | $4.00 | $5.00 | 20% lower |
| Output / 1M tokens | $20.00 | $25.00 | 20% lower |
| Cache reads / 1M tokens | $0.20 | $0.50 | 60% lower |
| Cache writes / 1M tokens | $5.00 | $6.25 | 20% lower |
The 40% figure is not a 40% cut to every token price. Input and output list rates are 20% lower. Anthropic says typical token-billed workloads cost about 40% less overall because Opus 5.5 also uses fewer tokens per task, while cache reads are 60% cheaper.
Why cache pricing matters for long-running agents
Long-running coding and agent workflows repeatedly reuse context: repository state, instructions, tool results, documents and earlier conversation. Prompt caching can make that repeated context much cheaper to read. Anthropic says cache reads make up a large share of agentic and coding-work costs, so reducing the cache-read rate from $0.50 to $0.20 per million tokens can matter disproportionately for workflows that preserve and reuse large contexts.
A simple cost scenario
Suppose a workload uses 20M uncached input tokens, 100M cache-read tokens and 5M output tokens. Using the published list rates, Opus 5 would cost about $250: $100 input + $50 cache reads + $100 output. Opus 5.5 would cost about $200: $80 input + $20 cache reads + $100 output. That is a 20% reduction before accounting for any reduction in tokens or steps per completed task.
The metric that matters is cost per successful task
For agents, token price alone is incomplete. A model that needs fewer turns, tool calls or retries can be cheaper even when its per-token rate is high. Anthropic reports that Opus 5.5 uses fewer tokens per task and cites early customer evaluations showing fewer steps in some coding and knowledge-work workloads. Those results are useful signals, but they are vendor and early-customer evidence rather than a guarantee for every production workflow.
If you already run Opus 5 agents, the practical test is not whether Opus 5.5 is 20% or 40% cheaper in the abstract. Replay a representative task set and measure total input, cache reads, output, tool calls, retries and successful completions. The best routing decision comes from cost per accepted result, not cost per million tokens.
Where Opus 5.5 is most likely to change the economics
- Long-running coding agents that repeatedly reuse repository context.
- Multi-tool agents where fewer turns or retries can reduce both model and tool costs.
- Document and knowledge workflows with large reusable context that benefits from prompt caching.
- High-value professional tasks where a more expensive frontier model can still be economical if it reduces rework.
What to benchmark before switching
A practical migration check
| Measure | Why it matters |
|---|---|
| Successful task rate | Cheaper tokens do not help if more tasks need manual repair. |
| Total output tokens | Verbose output can dominate frontier-model cost. |
| Cache-read volume | Opus 5.5’s 60% lower cache-read price can materially change repeated-context workloads. |
| Turns and tool calls | Fewer steps can lower latency and external tool costs. |
| Cost per accepted result | Combines quality and spend into a production metric you can actually route on. |
Limits on the headline numbers
- Anthropic’s roughly 40% typical-workload saving is a vendor estimate, not a universal discount.
- The exact saving depends on token mix, caching behavior, task complexity and how many steps the model needs.
- Fast mode uses different pricing: Anthropic lists $8 per million input tokens and $40 per million output tokens.
- US-only inference carries a pricing multiplier, so deployment requirements can change the economics.
- Benchmark results and early customer evaluations should be validated against your own workload before a production migration.
Opus 5.5’s list input and output prices are 20% below Opus 5, while cache reads are 60% cheaper. Anthropic says the combination of lower rates and better token efficiency brings typical token-billed workloads to about 40% lower cost. For agent builders, the useful next step is to benchmark cost per successful task—especially on long-running, cache-heavy workflows.
Sources & useful resources
- Anthropic — Introducing Claude Opus 5.5— Official release, pricing, efficiency claims and benchmark methodology.
- Anthropic — Claude Opus— Official availability, pricing and use-case summary.
- Reuters — Anthropic unveils Claude Opus 5.5— Independent release context.