Claude Haiku 5.5 Explained
$0.10/$0.50 Pricing, Benchmarks vs Haiku 4.5, and a Migration Guide (October 2026)
Claude Haiku 5.5 (Oct 7, 2026): $0.10/$0.50 per 1M tokens, 1M context, about a tenth of Haiku 4.5. Covers the 100K 5x price tier and 10 breaking changes.
Claude Haiku 5.5 is Anthropic's smallest, fastest-tier model, released on October 7, 2026. It costs $0.10 per million input tokens and $0.50 per million output tokens, roughly one tenth of Haiku 4.5 ($1 / $5), and its context window grows from 200K to 1M tokens. Anthropic describes the average saving as about 75%. It suits high-volume, latency-sensitive work such as summarization, classification, routing, extraction, ticket triage and subagents, while complex agentic coding still belongs to Sonnet 5.5 and Opus 5.5. Two caveats matter: the new tokenizer produces about 30% more tokens for the same text, and prompts over 100K tokens are billed at five times the rate.
This article covers the specs, benchmarks, how to estimate real cost, where Haiku 5.5 fits in the lineup, and the ten breaking changes when migrating from Haiku 4.5. All benchmark figures are vendor-reported by Anthropic and have not been independently reproduced. For the larger models, see the Claude Sonnet 5.5 guide and the Claude Opus 5.5 guide.
Claude Haiku 5.5 specifications
| Item | Details |
|---|---|
| Release date | October 7, 2026 |
| Model ID | claude-haiku-5-5 (Claude API / Google Cloud Vertex / Microsoft Foundry / Claude Platform on AWS). No date suffix |
| Amazon Bedrock ID | anthropic.claude-haiku-5-5 |
| Context window | 1M tokens (Haiku 4.5: 200K) |
| Max output | 128K tokens (up to 300K on the Batch API with a beta header) |
| Input | Text and images |
| Knowledge cutoff | June 2026 |
| Price (prompt up to 100K) | $0.10 input / $0.50 output per 1M tokens |
| Price (prompt over 100K) | $0.50 input / $2.50 output |
| Cache (up to 100K) | Read $0.01 / 5-minute write $0.125 |
| Cache (over 100K) | Read $0.05 / 5-minute write $0.625 |
| Batch API | 50% discount |
| Thinking | Adaptive thinking supported. First Haiku with adjustable effort (default medium) |
| Priority Tier | Not supported |
Adjustable effort is new for the Haiku line. At low, the model can skip thinking on simple requests, so classification and extraction can run at the lowest cost and latency. The default is medium, so thinking will occur unless you lower it.
Benchmarks: Haiku 5.5 vs Haiku 4.5 vs Sonnet 5.5
These are Anthropic's published figures (vendor-reported), listed as Haiku 5.5 / Haiku 4.5 / Sonnet 5.5.
| Benchmark | Haiku 5.5 | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|
| Terminal-Bench 4.0 | 39.2% | 0.0% | 70.6% |
| FrontierCode 1.1 (Main) | 46.4% | n/a | 52.1% (xhigh) |
| OSWorld 2.1 (offline subset) | 72.4% | 15.7% | 83.9% |
| Humanity's Last Exam (no tools) | 45.9% | 10.2% | 56.9% |
| GDPval-AA v2.1 (Elo) | 1620 | 735 | 1840 |
The jump from Haiku 4.5 is very large: Terminal-Bench 4.0 goes from 0.0% to 39.2% and OSWorld 2.1 from 15.7% to 72.4%. Read the OSWorld number carefully, though. The 72.4% includes partial credit; the rate at which all checkpoints are completed is 37.1%. It does not mean 70% of browser tasks succeed end to end. The Haiku 4.5 figures of 0.0% and 15.7% may also reflect different evaluation conditions, so it is safer to ask whether Haiku 5.5 has entered a usable range than to focus on the multiples.
A gap to Sonnet 5.5 remains. Terminal-Bench 4.0 is 39.2% versus 70.6%, so for long terminal-driven agentic coding, Haiku 5.5 is not a drop-in replacement for Sonnet 5.5. FrontierCode 1.1 looks closer (46.4% vs 52.1%), but the Sonnet figure uses xhigh effort, so the two are not directly comparable. On the third-party side, Artificial Analysis reports an Intelligence Index of 38 at high effort and 43 at max effort.
Pricing and real cost: the 100K boundary and the tokenizer
The unit price is one tenth of Haiku 4.5, but your actual bill depends on three things.
- The 100K-token boundary: with a prompt of 100K or less you pay $0.10 / $0.50; above 100K the whole request is billed at $0.50 / $2.50 (5x). Designs that fill the 1M window undo the savings
- The new tokenizer: the same text becomes about 30% more tokens than on Haiku 4.5, which narrows the saving
- Caching: reads cost $0.01 ($0.05 above 100K), so cache design dominates cost for workloads with long system prompts or tool definitions
A worked example (my own arithmetic, not an official figure). Assume 10M input and 2M output tokens per month measured on Haiku 4.5, with every request at or under 100K.
- Haiku 4.5: 10M x $1 + 2M x $5 = $20
- Haiku 5.5 (same token count): 10M x $0.10 + 2M x $0.50 = $2
- Haiku 5.5 with the ~30% tokenizer increase: 13M x $0.10 + 2.6M x $0.50 = about $2.60 (roughly 87% below Haiku 4.5)
- Same volume with every prompt over 100K: 13M x $0.50 + 2.6M x $2.50 = about $13 (roughly 35% below Haiku 4.5)
- For reference, the same volume on Sonnet 5.5 ($2 / $10) is $40
Workloads that fit under 100K see large savings, while pipelines that push whole documents into the prompt see much less. Always recount tokens on the new model and check which price tier each request falls into. The Batch API halves the price again for classification or extraction that can run asynchronously.
Lineup comparison: Opus 5.5 / Sonnet 5.5 / Haiku 5.5 / Haiku 4.5
| Model | Input / Output ($/1M) | Context | Role |
|---|---|---|---|
| Claude Opus 5.5 | $4 / $20 | 1M | Planning, review, open-ended hard work |
| Claude Sonnet 5.5 | $2 / $10 | 1M | Default for agentic coding |
| Claude Haiku 5.5 | $0.10 / $0.50 ($0.50 / $2.50 above 100K) | 1M | High volume, low latency, subagents |
| Claude Haiku 4.5 | $1 / $5 | 200K | Previous generation (migration source) |

Anthropic says Sonnet 5.5 and Opus 5.5 remain better for complex agentic coding. A typical split is Opus 5.5 for planning and review, Sonnet 5.5 for implementation, and Haiku 5.5 as subagent or router. This pairs naturally with the multi-agent patterns in the Claude Code and Agent SDK guide.
Which tasks to route to Haiku 5.5
Anthropic recommends Haiku 5.5 for the following.
- Summarization and conversation compaction
- Classification and routing (triaging inquiries, choosing which model handles a request)
- Structured data extraction
- Database query generation
- Support ticket triage
- Subagents that return short reports (research and search roles)
- Latency-sensitive support chat
- Browser use
It is not recommended for complex agentic coding. Long autonomous sessions and design or debugging that builds on many decisions should go to Sonnet 5.5 or Opus 5.5, with Haiku 5.5 handling the surrounding high-frequency, lightweight steps. A useful test is whether a failure is cheap to retry. Classification and triage are cheap to redo, which is where Haiku's low price pays off.
When using it as a router, set effort to low and keep outputs short. Output tokens dominate, so returning only a label keeps the per-request cost tiny.
Migrating from Haiku 4.5: ten breaking changes
The first step is changing the model ID from claude-haiku-4-5 to claude-haiku-5-5, but swapping the string alone will not work in several places. Switch in a test environment and fix errors in the order they appear.
- 1. Update the model ID: claude-haiku-5-5 (anthropic.claude-haiku-5-5 on Bedrock)
- 2. Recount tokens: the new tokenizer adds about 30%, so revisit max_tokens and cost estimates
- 3. Thinking configuration: thinking: {"type": "enabled", "budget_tokens": N} returns 400. Use {"type": "adaptive"} and set effort through output_config. Thinking tokens count toward max_tokens
- 4. Select content blocks by type, not position: the first block is not necessarily text (a thinking block may come first)
- 5. Sampling parameters: top_k, non-default temperature / top_p (or both together) return 400
- 6. Assistant prefill returns 400: you can no longer pin the start of the output; use instructions or structured outputs
- 7. Computer use tool: the old computer_20250124 returns 400; use computer_toolset_20260801
- 8. Thinking blocks are bound to the producing account: blocks from another account cannot be used
- 9. Keep conversations append-only: do not resend earlier thinking blocks after editing the system prompt, tools or earlier messages
- 10. Handle refusals: handle stop_reason: "refusal"; there is no server-side fallback
Items 3, 5 and 6 are the easiest to miss. A shared wrapper that always sends temperature, or code that steers JSON output with a { prefill, will start failing with 400s all at once. Item 4 is the classic breakage: code written as content[0].text fails as soon as a thinking block comes first.
Code example: adaptive thinking and effort
Here is a before-and-after in the Python SDK. Parameter placement can vary by SDK version, so check the official docs for the current syntax.
# Haiku 4.5 style (returns 400 on Haiku 5.5)
response = client.messages.create(
model="claude-haiku-4-5",
max_tokens=1024,
temperature=0.3,
thinking={"type": "enabled", "budget_tokens": 512},
messages=[
{"role": "user", "content": "Classify this inquiry: ..."},
{"role": "assistant", "content": "{"}, # prefill also returns 400
],
)
text = response.content[0].text # positional access is fragile too# Rewritten for Haiku 5.5
response = client.messages.create(
model="claude-haiku-5-5",
max_tokens=2048, # thinking tokens count toward this, so leave headroom
thinking={"type": "adaptive"},
output_config={"effort": "low"}, # low for simple work like classification
messages=[
{"role": "user", "content": "Classify this inquiry and return JSON only: ..."},
],
)
if response.stop_reason == "refusal":
handle_refusal(response) # no server-side fallback
else:
# pick the text block by type, not by position
text = next(b.text for b in response.content if b.type == "text")The rewrite drops temperature, replaces the prefill with an instruction, branches on stop_reason, and selects blocks by type. Start at low effort and raise it to medium (the default) or higher only for tasks where accuracy falls short.
Migration checklist
- Changed the model ID to claude-haiku-5-5 (anthropic.claude-haiku-5-5 on Bedrock)
- Recounted tokens on representative prompts and updated max_tokens and cost estimates
- Measured the share of requests over 100K tokens (5x price above that)
- Replaced thinking: enabled + budget_tokens with adaptive plus effort in output_config
- Removed top_k and any non-default temperature / top_p
- Removed assistant prefill, replacing it with instructions or structured outputs
- Changed response parsing from content[0] to selection by type
- Updated computer use to computer_toolset_20260801
- Kept conversation history append-only, with no stale thinking blocks resent after editing system or tools
- Added handling for stop_reason: "refusal"
- Compared low and medium effort on success rate, tokens and latency
Run 20 to 50 representative tasks at different effort levels and compare success rate and cost. If the success rate is flat, simply choosing the lower effort improves both cost and latency.
FAQ
How much does Claude Haiku 5.5 cost?
For prompts up to 100K tokens it is $0.10 input / $0.50 output per 1M tokens; above 100K it is $0.50 / $2.50. Cache reads are $0.01 ($0.05 above 100K) and the Batch API is 50% off. Anthropic says this is about 75% cheaper on average than Haiku 4.5 ($1 / $5).
Is it really 75% cheaper than Haiku 4.5?
The unit price is about one tenth, but the new tokenizer yields about 30% more tokens for the same text and prompts over 100K cost five times as much. Your real saving depends on those two factors, so recount tokens on the new model before estimating.
Can Haiku 5.5 replace Sonnet 5.5?
For lightweight tasks yes, but not for complex agentic coding. Terminal-Bench 4.0 is 39.2% for Haiku 5.5 and 70.6% for Sonnet 5.5 (both Anthropic-reported), and Anthropic itself says Sonnet 5.5 and Opus 5.5 remain better there.
Will my Haiku 4.5 code work after changing the model ID?
Possibly not. thinking enabled with budget_tokens, top_k, non-default temperature or top_p, assistant prefill and the old computer use tool all return 400. Switch in a test environment and fix errors as they appear.
What is Haiku 5.5 best suited for?
Summarization, compaction, classification, routing, extraction, database queries, ticket triage, subagents, latency-sensitive support chat and browser use. Complex agentic coding is better served by Sonnet 5.5 or Opus 5.5.
Feel free to contact us
Contact Us