AI models & agent economics

Claude Opus 5.5 cuts agent costs: what the new pricing changes

•Make Better Editorial

Anthropic says Claude Opus 5.5 costs about 40% less than Opus 5 on typical token-billed workloads. Here is what changed in token pricing, cache reads, and the economics of long-running agents.

Claude Opus 5.5 changes the economics of Anthropic’s top-tier model in two different ways. The easy-to-see change is a 20% cut to list input and output token prices. The more important claim for agent builders is Anthropic’s estimate that typical token-billed workloads cost about 40% less than Opus 5 because the new model also uses fewer tokens per task and makes cache reads much cheaper.

What actually got cheaper

Opus 5.5 vs Opus 5 list pricing

Cost itemOpus 5.5Opus 5Change
Input / 1M tokens$4.00$5.0020% lower
Output / 1M tokens$20.00$25.0020% lower
Cache reads / 1M tokens$0.20$0.5060% lower
Cache writes / 1M tokens$5.00$6.2520% lower
IMPORTANT DISTINCTION

The 40% figure is not a 40% cut to every token price. Input and output list rates are 20% lower. Anthropic says typical token-billed workloads cost about 40% less overall because Opus 5.5 also uses fewer tokens per task, while cache reads are 60% cheaper.

Why cache pricing matters for long-running agents

Long-running coding and agent workflows repeatedly reuse context: repository state, instructions, tool results, documents and earlier conversation. Prompt caching can make that repeated context much cheaper to read. Anthropic says cache reads make up a large share of agentic and coding-work costs, so reducing the cache-read rate from $0.50 to $0.20 per million tokens can matter disproportionately for workflows that preserve and reuse large contexts.

A simple cost scenario

SCENARIO, NOT A PRODUCT GUARANTEE

Suppose a workload uses 20M uncached input tokens, 100M cache-read tokens and 5M output tokens. Using the published list rates, Opus 5 would cost about $250: $100 input + $50 cache reads + $100 output. Opus 5.5 would cost about $200: $80 input + $20 cache reads + $100 output. That is a 20% reduction before accounting for any reduction in tokens or steps per completed task.

The metric that matters is cost per successful task

For agents, token price alone is incomplete. A model that needs fewer turns, tool calls or retries can be cheaper even when its per-token rate is high. Anthropic reports that Opus 5.5 uses fewer tokens per task and cites early customer evaluations showing fewer steps in some coding and knowledge-work workloads. Those results are useful signals, but they are vendor and early-customer evidence rather than a guarantee for every production workflow.

MAKE BETTER ANALYSIS

If you already run Opus 5 agents, the practical test is not whether Opus 5.5 is 20% or 40% cheaper in the abstract. Replay a representative task set and measure total input, cache reads, output, tool calls, retries and successful completions. The best routing decision comes from cost per accepted result, not cost per million tokens.

Where Opus 5.5 is most likely to change the economics

  • Long-running coding agents that repeatedly reuse repository context.
  • Multi-tool agents where fewer turns or retries can reduce both model and tool costs.
  • Document and knowledge workflows with large reusable context that benefits from prompt caching.
  • High-value professional tasks where a more expensive frontier model can still be economical if it reduces rework.

What to benchmark before switching

A practical migration check

MeasureWhy it matters
Successful task rateCheaper tokens do not help if more tasks need manual repair.
Total output tokensVerbose output can dominate frontier-model cost.
Cache-read volumeOpus 5.5’s 60% lower cache-read price can materially change repeated-context workloads.
Turns and tool callsFewer steps can lower latency and external tool costs.
Cost per accepted resultCombines quality and spend into a production metric you can actually route on.

Limits on the headline numbers

  • Anthropic’s roughly 40% typical-workload saving is a vendor estimate, not a universal discount.
  • The exact saving depends on token mix, caching behavior, task complexity and how many steps the model needs.
  • Fast mode uses different pricing: Anthropic lists $8 per million input tokens and $40 per million output tokens.
  • US-only inference carries a pricing multiplier, so deployment requirements can change the economics.
  • Benchmark results and early customer evaluations should be validated against your own workload before a production migration.
BOTTOM LINE

Opus 5.5’s list input and output prices are 20% below Opus 5, while cache reads are 60% cheaper. Anthropic says the combination of lower rates and better token efficiency brings typical token-billed workloads to about 40% lower cost. For agent builders, the useful next step is to benchmark cost per successful task—especially on long-running, cache-heavy workflows.

Sources & useful resources