Claude Haiku 5.5, which Anthropic released on October 7, 2026, costs $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. That is a tenth of Haiku 4.5's per-token rate. It also has a much larger context window and big gains on agentic benchmarks. It is not a drop-in swap, though. A new tokenizer, adaptive thinking and five breaking API changes mean your real saving and your migration effort will depend on your workload.
This guide covers what changed, what the independent numbers say, and the order to migrate in.
What Anthropic shipped
According to the official model overview, Claude Haiku 5.5 is aimed at "high-volume, latency-sensitive tasks such as classification, extraction, and routing." The core specs:
- Model ID:
claude-haiku-5-5on the Claude API, Google Cloud, Microsoft Foundry and Claude Platform on AWS. On Amazon Bedrock it isanthropic.claude-haiku-5-5. The ID has no date suffix and no separate alias. - Context window: 1M tokens, up from 200K on Haiku 4.5.
- Max output: 128K tokens, or up to 300K on the Message Batches API with the
output-300k-2026-03-24beta header. - Inputs: text and images. Output is text.
- Knowledge cutoff: June 2026.
- Lifecycle: Anthropic says the model will not be retired before October 7, 2027.
The announcement calls it Anthropic's fastest model at standard speed. It is also the first Haiku-class model with the effort parameter. Anthropic lists summaries, context compaction, database queries, classification, live customer support, browser use, and subagent work alongside Opus 5.5 or Sonnet 5.5 as target uses.
Pricing compared
| Per million tokens | Haiku 5.5 (≤100K prompt) | Haiku 5.5 (>100K prompt) | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|---|
| Input | $0.10 | $0.50 | $1.00 | $2.00 |
| Output | $0.50 | $2.50 | $5.00 | $10.00 |
| Cache read | $0.01 | $0.05 | $0.10 | $0.10 |
| 5-min cache write | $0.125 | $0.625 | $1.25 | $2.50 |
Sources: Anthropic's announcement and pricing in the model docs. The Batch API takes 50% off input and output. Anthropic also cut Sonnet 5.5's cache-read price from $0.20 to $0.10 on the same day.
Why the real saving is closer to 75% than 90%
The 90% figure is the per-token rate for short prompts. Anthropic's own number is that Haiku 5.5 is about 75% cheaper to run on average. Three things account for the gap.
- A new tokenizer. Haiku 5.5 uses the tokenizer introduced with Claude 4.7. According to the migration guide, the same text produces about 30% more tokens than on Haiku 4.5, and the exact increase depends on the content. Prompt sizes,
max_tokenslimits and cost estimates from Haiku 4.5 all need to be measured again. - The 100K price step. Prompts over 100,000 tokens cost five times as much per token. The 1M window is real, but long-context requests are priced very differently from short ones.
- Thinking tokens. Adaptive thinking is on by default and billed as output. The higher the effort level, the more the model thinks and the more you pay.
A worked example makes the trade-off concrete. Take a classification call that used 2,000 input and 300 output tokens on Haiku 4.5, which cost about $0.0035. On Haiku 5.5 the same text becomes roughly 2,600 input and 390 output tokens, about $0.00046, or 87% cheaper. If medium effort then adds 1,000 thinking tokens, the call costs about $0.00096. That is still roughly 73% cheaper than before, but noticeably less than the headline. At low effort, the model can skip thinking on simple requests altogether.
To estimate your own bill, count tokens with the new model ID rather than reusing old counts. Anthropic's token counting endpoint does this when model is set to claude-haiku-5-5. Then run the numbers through a spend model such as the LLM API Cost Calculator.
Benchmarks: vendor numbers vs. independent ones
Anthropic's announcement shows large gains over Haiku 4.5. Sonnet 5.5 still leads on every benchmark listed.
| Benchmark | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 (Elo) | 1620 | 735 | 1437 | 1840 |
| OSWorld 2.1 (offline subset) | 72.4% | 15.7% | 48.9% | 83.9% |
| Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 (Main) | 46.4% | — | 42.4% | 52.1% |
| Humanity's Last Exam (with tools) | 57.4% | 18.7% | — | 64.5% |
All figures are from Anthropic's announcement. Two early customers are quoted there. Box reported a score 11 points higher than Haiku 4.5 at about half the latency. AlphaSense reported 0.84 versus 0.76 on its Ask in Document evaluation.
Artificial Analysis's independent write-up mostly supports the direction of these results but adds useful context:
- Intelligence Index: Haiku 5.5 scores 43 at max effort and 38 at high. That puts it near Gemini 3.8 Flash (41) and GLM-5.3 Flash (42), slightly below Kimi K3 (44), above GPT-6 Luna (38), and 13 points behind Sonnet 5.5 at max (56).
- Verbosity: At max effort it used about 162K output tokens per Index task, roughly three times GPT-6 Luna's ~50K for a lower score. Artificial Analysis says it uses more output tokens than Opus 5.5 at max.
- Terminal-Bench 4.0: Artificial Analysis measured 33%, compared with Anthropic's 39.2%. Different harnesses produce different numbers, so treat the vendor figure as an upper bound.
- Knowledge vs. hallucination: It scored 36% accuracy on AA-Omniscience, below Gemini 3.8 Flash (55%) and GPT-6 Luna (44%). Its hallucination rate was lower, at 40% compared with 55% and 77%. In practice it answers fewer recall questions correctly but makes up answers less often.
- AutomationBench-AA: It scored 35% here, behind competitors at 53–60%. Artificial Analysis says the figure is probably understated because of a safety-refusal issue.
At medium effort, the API default, the Artificial Analysis model page lists an Index score of 34, about $0.05 per Index task, 155 output tokens per second, and a 13.4-second time to first token on Anthropic's API. Artificial Analysis notes that its cost figures do not yet include the price step above 100K tokens.
The takeaway: max effort makes Haiku competitive with other small models, but it spends a lot of tokens to get there. The per-token price does not tell you the per-task cost.
Picking an effort level
Haiku 5.5 supports all five levels: low, medium, high, xhigh and max. The effort docs confirm this, although some early coverage left out xhigh. The default is medium. Anthropic's guidance for this model:
| Effort | Use it for | Watch out for |
|---|---|---|
low |
Chat, short tool tasks, simple high-volume requests | In long agent prompts it is more likely to skip a search, stop early or skip a check |
medium (default) |
Most work, including agentic coding | Thinking still counts toward max_tokens |
high |
Knowledge work, longer agent tasks, strict instruction following | Higher output spend |
xhigh / max |
Only where your evals show a gain | At this point, compare against Sonnet 5.5 on cost and quality |
Two details matter in production. First, you can send thinking: {"type": "disabled"} at high effort or below, but it returns a 400 error at xhigh or max. Second, on the Claude API and Google Cloud you can change effort mid-conversation with a per-message output_config behind the mid-conversation-output-config-2026-07-01 beta header. This keeps the prompt cache, whereas changing the top-level effort between requests resets it.
Migration checklist from Haiku 4.5
The migration guide lists ten items. Here they are in the order most likely to break something:
- Swap the model ID for each platform (see above).
- Remove
temperature,top_pandtop_k. Any non-default value returns a 400 error. So doestop_p: 1, and so does sending bothtemperatureandtop_p. - Replace
budget_tokensthinking with{"type": "adaptive"}plusoutput_config.effort. Manual budgets now return a 400 error. - Remove assistant prefill. Requests must end with a user turn. For output formatting, use structured outputs, or tools with enum fields for classifiers. On Bedrock, which doesn't support structured outputs, use tools.
- Parse content blocks by
type, not position. A response can start with thinking blocks, and by default those come back with an emptythinkingfield. Setdisplay: "summarized"if you want the text. - Raise
max_tokens. Thinking counts against it, so a tight limit can returnstop_reason: "max_tokens"before any text appears. - Handle
stop_reason: "refusal". Safety classifiers can decline a request, and there is no server-side fallback model. - Keep conversations append-only when you send thinking blocks back. Editing
system,toolsor earlier messages invalidates them. - Replay stored conversations through the account that created them. Thinking blocks only work in the originating account or a linked one. Multi-tenant products that share a conversation store should check this.
- Computer use: move from
computer_20250124to thecomputer_toolset_20260801toolset on the Claude API and Google Cloud. A new browser toolset (browser_toolset_20260801) is also available there.
If you use Claude Code, /claude-api migrate this project to claude-haiku-5-5 applies the ID swap, parameter fixes, prefill replacement and effort calibration, then gives you a checklist to verify by hand. Claude Managed Agents users only need to change the model name.
One capacity caveat
Priority Tier is not supported on Haiku 5.5. As The Decoder points out, teams with a Priority Tier commitment on Haiku 4.5 have to plan capacity separately. Haiku 4.5 is still listed at $1/$5 with no announced retirement date, so you can migrate gradually.
Where Haiku 5.5 fits and where it doesn't
Good fits:
- Classification, routing and extraction at scale. At
lowormediumeffort with short prompts, the 10x lower rate goes furthest here. In DataCamp's small invoice-triage test, both Haiku models scored 24/24, but Haiku 5.5 cost $0.0098 compared with $0.25 for Haiku 4.5. The author notes the test was too easy to separate them on accuracy. - Subagents under Opus or Sonnet. Search, summarization and compaction steps were often where Haiku 4.5 hit its limits.
- Browser and computer-use automation. The OSWorld result moved from 15.7% to 72.4% on Anthropic's numbers, so the model can now handle work it previously could not.
Weaker fits:
- Complex agentic coding. Anthropic itself recommends Sonnet 5.5 or Opus 5.5 here. Sonnet roughly doubles the Terminal-Bench score.
- Long-document workloads that stay above 100K tokens. At $0.50/$2.50 the gap with Sonnet 5.5 ($2/$10) is only 4x, and the extra 30% of tokens narrows it further.
- Fact-recall without retrieval. Artificial Analysis's Omniscience results suggest you should ground the model with your own data.
If you are building on Claude and comparing it with other models, the AI coding assistants and AI customer support tools categories show how other products split work across models.
What to do this week
- Recount your top 20 prompts with
model: "claude-haiku-5-5"and see how many cross 100K tokens after the tokenizer change. - Run an effort sweep (
low,medium,high) on your existing evals and compare cost per successful task, not per token. - Ship behind a flag. Route 5–10% of traffic to Haiku 5.5, and log
refusalstop reasons andmax_tokenscut-offs separately. - Watch for updates. Artificial Analysis has said cost-per-task numbers that include the 100K price step are coming. Check Anthropic's model deprecations page for a Haiku 4.5 retirement date before you commit to a gradual migration.
Frequently asked questions
How much does Claude Haiku 5.5 cost?
For prompts up to 100,000 tokens it costs $0.10 per million input tokens and $0.50 per million output tokens. Above 100,000 tokens the rate goes up to $0.50 input and $2.50 output. Cache reads start at $0.01 per million tokens, and the Batch API takes 50% off input and output.
Is Claude Haiku 5.5 really 90% cheaper than Haiku 4.5?
The per-token rate is 90% lower for prompts under 100K tokens. Anthropic puts the average saving at about 75%, because the new tokenizer turns the same text into roughly 30% more tokens, prompts over 100K cost more, and adaptive thinking adds output tokens. Measure the savings on your own traffic instead of assuming 90%.
Is Claude Haiku 5.5 a drop-in replacement for Haiku 4.5?
No. Requests that use budget_tokens thinking, non-default temperature, top_p or top_k, assistant prefill, or the old computer_20250124 tool return 400 errors. Code also has to handle thinking blocks at the start of responses and the new refusal stop reason. Anthropic's migration guide lists ten items, and Claude Code's /claude-api migrate command can apply most of them.
What is the context window of Claude Haiku 5.5?
Anthropic's docs list a 1M-token context window and 128K max output tokens. On the Message Batches API, output can go up to 300K tokens with a beta header. Haiku 4.5 had 200K context and 64K output.
Should I use Claude Haiku 5.5 or Sonnet 5.5?
Use Haiku 5.5 for high-volume, well-scoped work such as classification, extraction, routing, support replies and subagents. Sonnet 5.5 scores higher on every benchmark Anthropic published, most of all on agentic coding (70.6% vs 39.2% on Terminal-Bench 4.0), and Anthropic still recommends Sonnet or Opus for complex agentic coding.
Sources
- Anthropic: Claude Haiku 5.5 announcement— anthropic.com
- Claude Platform Docs: Claude Haiku 5.5 overview— platform.claude.com
- Claude Platform Docs: Claude Haiku 5.5 migration guide— platform.claude.com
- Claude Platform Docs: Effort parameter— platform.claude.com
- Artificial Analysis: Anthropic has released Claude Haiku 5.5— artificialanalysis.ai
- Artificial Analysis: Claude Haiku 5.5 (medium) model page— artificialanalysis.ai
- DataCamp: Claude Haiku 5.5 features, benchmarks and pricing— datacamp.com
- The Decoder (mixed-news): Haiku 5.5 replaces Haiku 4.5, but Priority Tier does not come with it— mixed-news.com
