AnthropicClaudeLLM APIsAI pricingDeveloper tools

Claude Haiku 5.5: Pricing, Benchmarks and Migrating From Haiku 4.5

person
LaunchBoosts Research Desk·AI-assisted research
8 min read

Researched and drafted with AI assistance from the 8 public sources listed at the end of this article, then published after automated editorial checks. Spotted an error? Tell us at support@launchboosts.com.

Claude Haiku 5.5: Pricing, Benchmarks and Migrating From Haiku 4.5 — LaunchBoosts

Claude Haiku 5.5, which Anthropic released on October 7, 2026, costs $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. That is a tenth of Haiku 4.5's per-token rate. It also has a much larger context window and big gains on agentic benchmarks. It is not a drop-in swap, though. A new tokenizer, adaptive thinking and five breaking API changes mean your real saving and your migration effort will depend on your workload.

This guide covers what changed, what the independent numbers say, and the order to migrate in.

What Anthropic shipped

According to the official model overview, Claude Haiku 5.5 is aimed at "high-volume, latency-sensitive tasks such as classification, extraction, and routing." The core specs:

  • Model ID: claude-haiku-5-5 on the Claude API, Google Cloud, Microsoft Foundry and Claude Platform on AWS. On Amazon Bedrock it is anthropic.claude-haiku-5-5. The ID has no date suffix and no separate alias.
  • Context window: 1M tokens, up from 200K on Haiku 4.5.
  • Max output: 128K tokens, or up to 300K on the Message Batches API with the output-300k-2026-03-24 beta header.
  • Inputs: text and images. Output is text.
  • Knowledge cutoff: June 2026.
  • Lifecycle: Anthropic says the model will not be retired before October 7, 2027.

The announcement calls it Anthropic's fastest model at standard speed. It is also the first Haiku-class model with the effort parameter. Anthropic lists summaries, context compaction, database queries, classification, live customer support, browser use, and subagent work alongside Opus 5.5 or Sonnet 5.5 as target uses.

Pricing compared

Per million tokens Haiku 5.5 (≤100K prompt) Haiku 5.5 (>100K prompt) Haiku 4.5 Sonnet 5.5
Input $0.10 $0.50 $1.00 $2.00
Output $0.50 $2.50 $5.00 $10.00
Cache read $0.01 $0.05 $0.10 $0.10
5-min cache write $0.125 $0.625 $1.25 $2.50

Sources: Anthropic's announcement and pricing in the model docs. The Batch API takes 50% off input and output. Anthropic also cut Sonnet 5.5's cache-read price from $0.20 to $0.10 on the same day.

Why the real saving is closer to 75% than 90%

The 90% figure is the per-token rate for short prompts. Anthropic's own number is that Haiku 5.5 is about 75% cheaper to run on average. Three things account for the gap.

  1. A new tokenizer. Haiku 5.5 uses the tokenizer introduced with Claude 4.7. According to the migration guide, the same text produces about 30% more tokens than on Haiku 4.5, and the exact increase depends on the content. Prompt sizes, max_tokens limits and cost estimates from Haiku 4.5 all need to be measured again.
  2. The 100K price step. Prompts over 100,000 tokens cost five times as much per token. The 1M window is real, but long-context requests are priced very differently from short ones.
  3. Thinking tokens. Adaptive thinking is on by default and billed as output. The higher the effort level, the more the model thinks and the more you pay.

A worked example makes the trade-off concrete. Take a classification call that used 2,000 input and 300 output tokens on Haiku 4.5, which cost about $0.0035. On Haiku 5.5 the same text becomes roughly 2,600 input and 390 output tokens, about $0.00046, or 87% cheaper. If medium effort then adds 1,000 thinking tokens, the call costs about $0.00096. That is still roughly 73% cheaper than before, but noticeably less than the headline. At low effort, the model can skip thinking on simple requests altogether.

To estimate your own bill, count tokens with the new model ID rather than reusing old counts. Anthropic's token counting endpoint does this when model is set to claude-haiku-5-5. Then run the numbers through a spend model such as the LLM API Cost Calculator.

Benchmarks: vendor numbers vs. independent ones

Anthropic's announcement shows large gains over Haiku 4.5. Sonnet 5.5 still leads on every benchmark listed.

Benchmark Haiku 5.5 Haiku 4.5 GPT-6 Luna Sonnet 5.5
GDPval-AA v2.1 (Elo) 1620 735 1437 1840
OSWorld 2.1 (offline subset) 72.4% 15.7% 48.9% 83.9%
Terminal-Bench 4.0 39.2% 0.0% 16.4% 70.6%
FrontierCode 1.1 (Main) 46.4% — 42.4% 52.1%
Humanity's Last Exam (with tools) 57.4% 18.7% — 64.5%

All figures are from Anthropic's announcement. Two early customers are quoted there. Box reported a score 11 points higher than Haiku 4.5 at about half the latency. AlphaSense reported 0.84 versus 0.76 on its Ask in Document evaluation.

Artificial Analysis's independent write-up mostly supports the direction of these results but adds useful context:

  • Intelligence Index: Haiku 5.5 scores 43 at max effort and 38 at high. That puts it near Gemini 3.8 Flash (41) and GLM-5.3 Flash (42), slightly below Kimi K3 (44), above GPT-6 Luna (38), and 13 points behind Sonnet 5.5 at max (56).
  • Verbosity: At max effort it used about 162K output tokens per Index task, roughly three times GPT-6 Luna's ~50K for a lower score. Artificial Analysis says it uses more output tokens than Opus 5.5 at max.
  • Terminal-Bench 4.0: Artificial Analysis measured 33%, compared with Anthropic's 39.2%. Different harnesses produce different numbers, so treat the vendor figure as an upper bound.
  • Knowledge vs. hallucination: It scored 36% accuracy on AA-Omniscience, below Gemini 3.8 Flash (55%) and GPT-6 Luna (44%). Its hallucination rate was lower, at 40% compared with 55% and 77%. In practice it answers fewer recall questions correctly but makes up answers less often.
  • AutomationBench-AA: It scored 35% here, behind competitors at 53–60%. Artificial Analysis says the figure is probably understated because of a safety-refusal issue.

At medium effort, the API default, the Artificial Analysis model page lists an Index score of 34, about $0.05 per Index task, 155 output tokens per second, and a 13.4-second time to first token on Anthropic's API. Artificial Analysis notes that its cost figures do not yet include the price step above 100K tokens.

The takeaway: max effort makes Haiku competitive with other small models, but it spends a lot of tokens to get there. The per-token price does not tell you the per-task cost.

Picking an effort level

Haiku 5.5 supports all five levels: low, medium, high, xhigh and max. The effort docs confirm this, although some early coverage left out xhigh. The default is medium. Anthropic's guidance for this model:

Effort Use it for Watch out for
low Chat, short tool tasks, simple high-volume requests In long agent prompts it is more likely to skip a search, stop early or skip a check
medium (default) Most work, including agentic coding Thinking still counts toward max_tokens
high Knowledge work, longer agent tasks, strict instruction following Higher output spend
xhigh / max Only where your evals show a gain At this point, compare against Sonnet 5.5 on cost and quality

Two details matter in production. First, you can send thinking: {"type": "disabled"} at high effort or below, but it returns a 400 error at xhigh or max. Second, on the Claude API and Google Cloud you can change effort mid-conversation with a per-message output_config behind the mid-conversation-output-config-2026-07-01 beta header. This keeps the prompt cache, whereas changing the top-level effort between requests resets it.

Migration checklist from Haiku 4.5

The migration guide lists ten items. Here they are in the order most likely to break something:

  1. Swap the model ID for each platform (see above).
  2. Remove temperature, top_p and top_k. Any non-default value returns a 400 error. So does top_p: 1, and so does sending both temperature and top_p.
  3. Replace budget_tokens thinking with {"type": "adaptive"} plus output_config.effort. Manual budgets now return a 400 error.
  4. Remove assistant prefill. Requests must end with a user turn. For output formatting, use structured outputs, or tools with enum fields for classifiers. On Bedrock, which doesn't support structured outputs, use tools.
  5. Parse content blocks by type, not position. A response can start with thinking blocks, and by default those come back with an empty thinking field. Set display: "summarized" if you want the text.
  6. Raise max_tokens. Thinking counts against it, so a tight limit can return stop_reason: "max_tokens" before any text appears.
  7. Handle stop_reason: "refusal". Safety classifiers can decline a request, and there is no server-side fallback model.
  8. Keep conversations append-only when you send thinking blocks back. Editing system, tools or earlier messages invalidates them.
  9. Replay stored conversations through the account that created them. Thinking blocks only work in the originating account or a linked one. Multi-tenant products that share a conversation store should check this.
  10. Computer use: move from computer_20250124 to the computer_toolset_20260801 toolset on the Claude API and Google Cloud. A new browser toolset (browser_toolset_20260801) is also available there.

If you use Claude Code, /claude-api migrate this project to claude-haiku-5-5 applies the ID swap, parameter fixes, prefill replacement and effort calibration, then gives you a checklist to verify by hand. Claude Managed Agents users only need to change the model name.

One capacity caveat

Priority Tier is not supported on Haiku 5.5. As The Decoder points out, teams with a Priority Tier commitment on Haiku 4.5 have to plan capacity separately. Haiku 4.5 is still listed at $1/$5 with no announced retirement date, so you can migrate gradually.

Where Haiku 5.5 fits and where it doesn't

Good fits:

  • Classification, routing and extraction at scale. At low or medium effort with short prompts, the 10x lower rate goes furthest here. In DataCamp's small invoice-triage test, both Haiku models scored 24/24, but Haiku 5.5 cost $0.0098 compared with $0.25 for Haiku 4.5. The author notes the test was too easy to separate them on accuracy.
  • Subagents under Opus or Sonnet. Search, summarization and compaction steps were often where Haiku 4.5 hit its limits.
  • Browser and computer-use automation. The OSWorld result moved from 15.7% to 72.4% on Anthropic's numbers, so the model can now handle work it previously could not.

Weaker fits:

  • Complex agentic coding. Anthropic itself recommends Sonnet 5.5 or Opus 5.5 here. Sonnet roughly doubles the Terminal-Bench score.
  • Long-document workloads that stay above 100K tokens. At $0.50/$2.50 the gap with Sonnet 5.5 ($2/$10) is only 4x, and the extra 30% of tokens narrows it further.
  • Fact-recall without retrieval. Artificial Analysis's Omniscience results suggest you should ground the model with your own data.

If you are building on Claude and comparing it with other models, the AI coding assistants and AI customer support tools categories show how other products split work across models.

What to do this week

  1. Recount your top 20 prompts with model: "claude-haiku-5-5" and see how many cross 100K tokens after the tokenizer change.
  2. Run an effort sweep (low, medium, high) on your existing evals and compare cost per successful task, not per token.
  3. Ship behind a flag. Route 5–10% of traffic to Haiku 5.5, and log refusal stop reasons and max_tokens cut-offs separately.
  4. Watch for updates. Artificial Analysis has said cost-per-task numbers that include the 100K price step are coming. Check Anthropic's model deprecations page for a Haiku 4.5 retirement date before you commit to a gradual migration.

Frequently asked questions

How much does Claude Haiku 5.5 cost?

For prompts up to 100,000 tokens it costs $0.10 per million input tokens and $0.50 per million output tokens. Above 100,000 tokens the rate goes up to $0.50 input and $2.50 output. Cache reads start at $0.01 per million tokens, and the Batch API takes 50% off input and output.

Is Claude Haiku 5.5 really 90% cheaper than Haiku 4.5?

The per-token rate is 90% lower for prompts under 100K tokens. Anthropic puts the average saving at about 75%, because the new tokenizer turns the same text into roughly 30% more tokens, prompts over 100K cost more, and adaptive thinking adds output tokens. Measure the savings on your own traffic instead of assuming 90%.

Is Claude Haiku 5.5 a drop-in replacement for Haiku 4.5?

No. Requests that use budget_tokens thinking, non-default temperature, top_p or top_k, assistant prefill, or the old computer_20250124 tool return 400 errors. Code also has to handle thinking blocks at the start of responses and the new refusal stop reason. Anthropic's migration guide lists ten items, and Claude Code's /claude-api migrate command can apply most of them.

What is the context window of Claude Haiku 5.5?

Anthropic's docs list a 1M-token context window and 128K max output tokens. On the Message Batches API, output can go up to 300K tokens with a beta header. Haiku 4.5 had 200K context and 64K output.

Should I use Claude Haiku 5.5 or Sonnet 5.5?

Use Haiku 5.5 for high-volume, well-scoped work such as classification, extraction, routing, support replies and subagents. Sonnet 5.5 scores higher on every benchmark Anthropic published, most of all on agentic coding (70.6% vs 39.2% on Terminal-Bench 4.0), and Anthropic still recommends Sonnet or Opus for complex agentic coding.

Sources

  1. Anthropic: Claude Haiku 5.5 announcement— anthropic.com
  2. Claude Platform Docs: Claude Haiku 5.5 overview— platform.claude.com
  3. Claude Platform Docs: Claude Haiku 5.5 migration guide— platform.claude.com
  4. Claude Platform Docs: Effort parameter— platform.claude.com
  5. Artificial Analysis: Anthropic has released Claude Haiku 5.5— artificialanalysis.ai
  6. Artificial Analysis: Claude Haiku 5.5 (medium) model page— artificialanalysis.ai
  7. DataCamp: Claude Haiku 5.5 features, benchmarks and pricing— datacamp.com
  8. The Decoder (mixed-news): Haiku 5.5 replaces Haiku 4.5, but Priority Tier does not come with it— mixed-news.com
person

LaunchBoosts Research Desk

AI-assisted research

Explainers on software and AI industry trends, drafted with AI assistance from the public sources cited in each article and published after automated editorial checks for length, independent sourcing and originality.