Skip to main content
株式会社オブライト
AI2026-09-2912 min read

Claude Sonnet 5.5: $2/$10 Pricing, Benchmarks & Migration

Claude Sonnet 5.5 (Sep 28, 2026): $2/$10 per 1M tokens like Sonnet 5, 1M context, Terminal-Bench 4.0 at 70.6%. Sep 2026: benchmarks, 5 breaking changes.


Claude Sonnet 5.5 is Anthropic's mid-tier model, released on September 28, 2026. It costs $2 per million input tokens and $10 per million output tokens, the same as Sonnet 5, with a 1M-token context window and up to 128K output tokens. By Anthropic's published numbers, Terminal-Bench 4.0 rises from 10.3% to 70.6%, CursorBench 4.0 from 34.1% to 55.5%, and OSWorld 2.1 from 57.0% to 80.1%. Output is more than 30% faster, and cost per task is said to drop by up to 30% because the model makes fewer tool calls. Teams running agentic coding or computer use on Sonnet 5 have the most to gain, but the API has five breaking changes, so swapping the model string alone may not work.

This article lays out the specs, pricing, and benchmarks, then walks through the migration from Sonnet 5 for developers. All figures come from Anthropic's official page and documentation and have not been independently reproduced. It also covers how to choose between Sonnet 5.5 and the higher tiers, Claude Opus 5.5 and Claude Fable 5.1.

Claude Sonnet 5.5 Specs at a Glance

ItemDetails
Release dateSeptember 28, 2026
Model IDclaude-sonnet-5-5 (Claude API / Google Cloud / Microsoft Foundry / Claude Platform on AWS)
Amazon Bedrock IDanthropic.claude-sonnet-5-5
AvailabilityThe platforms above, the Claude apps, and Claude Code
Context window1M tokens
Max output128K tokens (up to 300K on the Batch API with the output-300k-2026-03-24 beta header)
ModalitiesText and images in, text out
Knowledge cutoffJune 2026
Pricing$2 input / $10 output per 1M tokens
Caching5-minute write $2.50 / 1-hour write $4 / read $0.20
Batch API50% discount
ThinkingAdaptive thinking on by default; default effort is high
TokenizerSame as Sonnet 5 (identical token counts for identical input)
Support windowNo retirement before September 28, 2027

The official documentation states the prices are the same as Claude Sonnet 5, and the tokenizer is unchanged, so the price of a given request does not change for the same prompt. What changes is how many tokens and tool calls it takes to finish the same job.

Benchmarks: Sonnet 5.5 vs Sonnet 5

Here is Anthropic's published comparison between Sonnet 5.5 and Sonnet 5.

BenchmarkSonnet 5.5Sonnet 5Delta
Terminal-Bench 4.070.6%10.3%+60.3pt
FrontierCode 1.1 (Main, Max effort)46.2%42.4%+3.8pt
CursorBench 4.055.5%34.1%+21.4pt
GDPval-AA v2.118441449+395
AA-Briefcase v1.118111359+452
OSWorld 2.180.1%57.0%+23.1pt

The Terminal-Bench 4.0 jump is the standout. On tasks that require completing long jobs while operating a terminal, the score leaps from 10.3% to 70.6%. The gap is so large that the Sonnet 5 figure should be read with possible differences in evaluation conditions in mind, and you should not assume this multiple of improvement until you have measured your own workload. FrontierCode 1.1, by contrast, moves only +3.8pt, so not every coding metric improved by the same margin.

On CursorBench 4.0, Sonnet 5.5 is reported to be within 2 points of the higher-tier Opus 5.5. Anthropic itself positions Opus 5.5 as clearly stronger at complex, open-ended work requiring sustained judgment, so a close benchmark score does not automatically mean it is a drop-in replacement. The +23.1pt on OSWorld 2.1 matters for computer-use agents, but it comes with a tool definition change, described below.

Pricing and Cost per Task

Because token prices are identical to Sonnet 5, any savings come from consumption per task, not from the unit price. Anthropic says faster execution and fewer tool calls can lower cost per task by up to 30%. That is a ceiling, not a guarantee.

Here is a monthly estimate for a workload that uses 10 million input tokens and 2 million output tokens.

- Standard: 10M input x $2 = $20, plus 2M output x $10 = $20, for a total of $40
- Batch API (50% off): $20 total (only for workloads that can run asynchronously)
- If 80% of input (8M tokens) is served from cache reads: 2M uncached x $2 = $4, plus 8M cache reads x $0.20 = $1.60, plus $20 output, for $25.60 total (a simplified calculation that excludes cache write costs)
- If cost per task falls by the 30% ceiling: the standard $40 becomes roughly $28 (a best-case assumption)

Since a cache read at $0.20 is one tenth of the input price, cache design is the biggest cost lever for agents with long system prompts or tool definitions. The minimum cacheable prompt drops from 1,024 tokens on Sonnet 5 to 512 tokens on Sonnet 5.5, so shorter prompts that previously could not be cached now qualify.

Lineup Comparison: Fable 5.1 / Opus 5.5 / Sonnet 5.5 / Haiku 4.5

ModelInput / Output ($/1M)SpeedDefault effortContext / Max output
Claude Fable 5.1$10 / $50Slowerhigh1M / 128K
Claude Opus 5.5$4 / $20Moderatemedium1M / 128K
Claude Sonnet 5.5$2 / $10Fasthigh1M / 128K
Claude Haiku 4.5$1 / $5Fastestn/a200K / 64K

Use the following as a rule of thumb. For details on each model, see the Fable 5.1 explainer and the Opus 5.5 explainer.

- Start with Sonnet 5.5: the default candidate for coding assistance, code generation, computer use, and internal knowledge search, where volume and speed both matter. It costs half of Opus 5.5 and is reported to be within 2 points on CursorBench
- Move up to Opus 5.5: long-running autonomous work and complex, ambiguous tasks that need accumulated judgment. It is the next step when Sonnet 5.5 keeps failing
- Reserve Fable 5.1 for the hardest work: it costs five times Sonnet 5.5 and is slower, so limit it to jobs that truly need top-tier capability
- Haiku 4.5 for high volume and low latency: classification, extraction, and other jobs that prioritize the fastest and cheapest option. Note its 200K context and 64K max output

In practice, making Sonnet 5.5 the default and routing only high-failure tasks to a higher tier is a cost-efficient setup.

Migrating from Sonnet 5: Five Breaking Changes

Migration map from Claude Sonnet 5 to Sonnet 5.5: the model ID becomes claude-sonnet-5-5 at the same $2 input / $10 output price, and five breaking changes (thinking disabled, tool_choice any/tool, editing history, computer_20251124, old advisor models) are replaced by between_tools, auto plus strict, append-only history, computer_toolset_20260801, and Opus 5/5.5, Fable, Mythos or Sonnet 5.5.

The first step is changing the model string from claude-sonnet-5 to claude-sonnet-5-5. However, Sonnet 5.5 has five changes that make existing code return 400 errors. The quickest path is to switch only the model string in a test environment and fix each error as it appears.

1. thinking: {"type": "disabled"} and manual budget_tokens return 400. You can no longer turn thinking off entirely. Use thinking: {"type": "between_tools"} instead: it is the lowest setting, disabling up-front thinking so the model reasons only between tool calls. It works only with effort set to low, medium, or high; combining it with xhigh or max returns 400.

2. Forced tool use is not supported. Setting tool_choice to {"type": "any"} or {"type": "tool", ...} returns a 400 error saying types "tool" and "any" are not supported for this model. Switch to auto and either add strict: true (strict tool use) to your tool definitions or use structured outputs to guarantee the schema. Designs that force a specific tool call, such as extraction pipelines, need a second look.

# Sonnet 5 style (returns 400 on Sonnet 5.5)
response = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=2048,
    thinking={"type": "disabled"},
    tool_choice={"type": "tool", "name": "extract_invoice"},
    tools=tools,
    messages=messages,
)
# Rewritten for Sonnet 5.5
response = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=2048,
    thinking={"type": "between_tools"},  # lowest setting; effort must be low / medium / high
    tool_choice={"type": "auto"},
    tools=[{**t, "strict": True} for t in tools],  # strict tool use guarantees the schema
    messages=messages,
)

3. Thinking blocks are tied to the model and the conversation. If you edit earlier system, tools, or messages and then replay a history that contains a Sonnet 5.5 thinking block, the request returns 400. This is enforced by default for accounts created on or after August 31, 2026. Thinking blocks are also bound to the account that produced them, and blocks from other accounts are dropped. Keep conversation history append-only, and use mid-conversation system messages (covered below) when you need to change instructions. Implementations that compress context by summarizing or deleting past messages need particular care.

4. The computer use tool definition changes. computer_20251124 is rejected on the Claude API and Google Cloud, so replace it with computer_toolset_20260801. Amazon Bedrock still accepts the old tool for now. This is required work if you want to benefit from the OSWorld 2.1 gains.

5. Advisor tool pairing is restricted. With Sonnet 5.5 as the executor, advisors of Opus 4.8, Opus 4.7, or Sonnet 5 are rejected. Choose a supported newer model as the advisor.

A change that does not raise an error also deserves attention. Text between tool calls is now returned as thinking progress-update blocks. At the default display setting (omitted) they are empty, so streaming UIs can appear to go quiet mid-response. Set thinking.display or use between_tools to handle this. Also, refusals now come back as HTTP 200 with stop_reason: "refusal", with stop_details categories of cyber, bio, frontier_llm, reasoning_extraction, or general_harms. Code that only watches for HTTP errors will miss them, so add handling.

Sampling parameters have changed too. Setting temperature, top_p, or top_k to a non-default value returns a 400 error. Few codebases may still carry this from Sonnet 5, but it is worth checking that no shared wrapper passes a temperature by default.

Migration Checklist

- Changed the model string to claude-sonnet-5-5 (anthropic.claude-sonnet-5-5 on Bedrock)
- Removed thinking: disabled and budget_tokens, replacing with between_tools where needed (never together with xhigh or max)
- Replaced tool_choice any / tool with auto plus strict: true, or with structured outputs
- Conversation history is append-only, and old thinking blocks are not replayed after editing system or tools
- Updated computer use to computer_toolset_20260801 (Claude API / Google Cloud)
- No advisors of Opus 4.8 / 4.7 / Sonnet 5 in the Advisor tool
- No non-default temperature / top_p / top_k
- Checked that streaming UIs do not appear stalled by thinking progress blocks
- Added handling for stop_reason: "refusal"
- Re-tuned effort (see the next sections)

New Features in Sonnet 5.5

- Per-message effort (beta): switch effort per request, even mid-conversation. Useful for dynamic tuning, such as low for simple questions and high only for hard steps
- Mid-conversation system messages: add instructions without rewriting history. This pairs with the append-only practice from breaking change 3
- Mid-conversation tool changes (beta): swap the available tools partway through a conversation
- Compaction on demand: with the compact-2026-09-04 beta header, compact a long conversation exactly when you choose
- Inline tool definitions (beta): with the inline-tools-2026-09-15 beta header, define tools inside a message
- Caching from 512 tokens: the minimum cacheable prompt falls from 1,024 tokens on Sonnet 5 to 512

Beta features may change, so confirm the latest official documentation before putting them into production.

Effort Tuning Tips

Effort levels on Sonnet 5.5 have been recalibrated from Sonnet 5. The same high does not necessarily behave the same, so re-run an effort sweep after migrating. The official guidance is as follows.

- Agentic coding: start at medium, and move to high for harder tasks
- Chat and latency-sensitive uses: medium or low
- Minimal thinking: combine low or similar with between_tools
- The default is high: if you specify nothing, the model thinks more deeply than Opus 5.5 does at its default of medium, so lower it while watching cost and latency

response = client.messages.create(
    model="claude-sonnet-5-5",
    max_tokens=4096,
    effort="medium",  # default is high; for agentic coding, start at medium
    messages=[{"role": "user", "content": "Fix the failing tests"}],
)

For evaluation, run 20 to 50 representative tasks at low, medium, and high effort and compare success rate, token count, and elapsed time. If success rate is flat, simply adopting the lower effort improves both cost and speed.

Frequently Asked Questions

How much does Claude Sonnet 5.5 cost?

It costs $2 per 1M input tokens and $10 per 1M output tokens, the same as Sonnet 5. Caching is $2.50 for a 5-minute write, $4 for a 1-hour write, and $0.20 for a read, and the Batch API is 50% off. Anthropic says cost per task can fall by up to 30% thanks to faster execution and fewer tool calls.

What is the difference between Sonnet 5.5 and Sonnet 5?

Pricing and the tokenizer are the same, while performance and behavior changed. Published figures show Terminal-Bench 4.0 rising from 10.3% to 70.6%, CursorBench 4.0 from 34.1% to 55.5%, and OSWorld 2.1 from 57.0% to 80.1%, with output more than 30% faster. The API also has five breaking changes, including disabled thinking and forced tool use.

Will my Sonnet 5 code work if I only change the model string?

It may not. Using thinking: disabled, tool_choice any or tool, the old computer use tool, or older advisor models will return 400 errors. The efficient approach is to switch only the model string in a test environment and fix the errors as they appear.

How do I turn thinking off?

You cannot fully disable it. Setting thinking type to between_tools is the lowest setting: it turns off up-front thinking so the model reasons only between tool calls. It works only with effort at low, medium, or high, and returns a 400 error with xhigh or max.

When should I choose Sonnet 5.5 over Opus 5.5?

For high-volume coding assistance and computer use where speed and cost matter, Sonnet 5.5 is the first candidate. It costs half of Opus 5.5 and is reported to be within 2 points on CursorBench. Anthropic says Opus 5.5 is stronger at long-running autonomous work and complex tasks that need sustained judgment.

Feel free to contact us

Contact Us