Claude Haiku 5.5 Guide: Pricing, Specs and Platforms
As of 2026-10-08, Claude Haiku 5.5 (model ID claude-haiku-5-5) has been out since 2026-10-07: a 1M context window and 128K maximum output, priced in two tiers by prompt length — input $0.10 / output $0.50 per million tokens for prompts up to 100,000 tokens, and input $0.50 / output $2.50 for prompts over 100,000 tokens. Anthropic calls it its most capable model tuned for high-volume and latency-sensitive work. This page covers the IDs, both price tiers, use cases, what changed from Haiku 4.5 and which platforms carry it.
Updated 2026-10-08
Four key numbers
Release date
The official model page reads "Released October 7, 2026", and the Claude API release notes list it the same day. Status: Active, with retirement no sooner than 2027-10-07.
Input / output, prompts ≤100K (per million tokens)
Rates for prompts up to 100,000 tokens; cache reads $0.01, 5-minute cache writes $0.125, 1-hour cache writes $0.20.
Input / output, prompts >100K
Once a prompt goes over 100,000 tokens, both input and output are billed at this tier: cache reads $0.05, 5-minute cache writes $0.625, 1-hour cache writes $1.
Context window / max output
Haiku 4.5 offers 200K / 64K. Input can be text and images; output is text.
What Haiku 5.5 is and what it is for
As of 2026-10-08, Claude Haiku 5.5 is the Haiku-class model Anthropic released on 2026-10-07; the official release notes call it "our most capable model tuned for high-volume and latency-sensitive work." The official model page lists classification, routing, extraction and subagent tasks as typical uses, and the launch post adds that it reliably handles quick, repetitive workloads such as summaries, compactions, database queries and classification requests, and pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding work. The same post says Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding. It is also the first Haiku with an effort setting: the default is medium, and thinking is on by default.
The vendor's own framing (released 2026-10-07)
The launch post says Haiku 5.5 costs much less than Haiku 4.5: "On average, it now costs around 75% less to run." A footnote explains the math: requests up to 100,000 tokens are 90% cheaper than on Haiku 4.5 and requests over 100,000 tokens are 50% cheaper, about 90% of Haiku 4.5 requests fell into the first group, and the figure already accounts for the new tokenizer using more tokens. That is Anthropic's own calculation. The pricing page adds that among Claude 4.6 and later models only Haiku 5.5 is priced by prompt length; the others bill the full 1M context at standard rates. For end users, Free, Pro, Max, Team and Enterprise users can select Haiku 5.5 on Claude.ai on web, iOS and Android, and it also runs in Claude Code.
The Haiku 5.5 timeline
Claude Haiku 4.5 is released (model ID claude-haiku-4-5-20251001) with a 200K context window and 64K maximum output.
Claude Haiku 5.5 ships as claude-haiku-5-5 with 1M / 128K and two price tiers by prompt length.
The official retirement date for Haiku 4.5 (claude-haiku-4-5-20251001) reads "Not sooner than October 15, 2026": a floor, not a shutdown date. Its status is still Active; watch the official deprecations page for the actual notice.
Confirmed vs not verified
Confirmed (verbatim in the official pages)
All of the following can be checked word for word on official pages: released 2026-10-07; model ID claude-haiku-5-5 (a fixed ID with no date suffix and no separate alias), anthropic.claude-haiku-5-5 on Amazon Bedrock; 1M context and 128K maximum output; two price tiers (up to 100K: input $0.10 / output $0.50; over 100K: input $0.50 / output $2.50); cache reads $0.01 / $0.05; the Batch API at 50% off input and output; default effort medium; about 30% more tokens for the same text; manual budget_tokens, non-default temperature / top_p / top_k and assistant prefill all return 400; retirement no sooner than 2027-10-07. The "around 75% less" figure and the use-case descriptions are the vendor's own statements.
Unverified and not published by the vendor
Four things are not covered by the official pages, so do not plan around guesses: first, the price list does not spell out how the 100,000-token threshold is counted (for example whether cache writes and cache hits are included), and this page does not infer it; compare your own usage data with the official price list. Second, "around 75% less" is Anthropic's own calculation based on Haiku 4.5's request mix, so your saving depends on how your own prompt lengths are distributed. Third, the benchmark scores and customer quotes in the launch post are vendor testing and customer statements, and this page does not reproduce them. Fourth, whether or when claude-haiku-5-5 comes to QCode has not been announced; check /models.
How it compares with Haiku 4.5 and Sonnet 5.5
vs Haiku 4.5
Haiku 4.5 (claude-haiku-4-5-20251001) is input $1 / output $5 with a 200K context window and 64K maximum output. Converted from the official price list: for prompts up to 100,000 tokens, Haiku 5.5 costs one tenth of Haiku 4.5 per token on input and output ($0.10 / $0.50 vs $1 / $5); over 100,000 tokens it costs half ($0.50 / $2.50 vs $1 / $5). The official docs also say the same text produces about 30% more tokens on Haiku 5.5, so recount before comparing costs. Haiku 5.5 adds effort and adaptive thinking, and raises the limits to 1M context and 128K output.
vs Sonnet 5.5
Sonnet 5.5 (claude-sonnet-5-5) has 1M / 128K at input $2 / output $10, with no prompt-length tiers. Converted from the official price list: for prompts up to 100,000 tokens Haiku 5.5's per-token rates are one twentieth of Sonnet 5.5's, and over 100,000 tokens they are one quarter. The official division of labor: Sonnet 5.5 and Opus 5.5 remain better choices for complex agentic coding, while Haiku 5.5 is best suited to narrowly scoped tasks that used to be cost-prohibitive with Claude, such as compaction, summarization or subagent work.
What to change when moving from Haiku 4.5
The official release notes put it plainly: "Code written for Claude Haiku 4.5 can break on Claude Haiku 5.5." The main points of the migration guide: first, set model to claude-haiku-5-5 (anthropic.claude-haiku-5-5 on Bedrock); second, recount your prompts with the new model and revisit max_tokens and cost estimates, because the same text is about 30% more tokens and thinking tokens count toward max_tokens; third, replace thinking enabled plus budget_tokens with adaptive and steer depth with effort (turning thinking off works only at low, medium and high; xhigh and max return 400); fourth, remove temperature, top_p and top_k, since non-default values return 400; fifth, end messages with a user turn, because an assistant prefill returns 400 even with thinking off; sixth, select content blocks by type, because a response can begin with thinking blocks.
On QCode
Whether claude-haiku-5-5 is listed on QCode is shown on /models. You can use claude-haiku-4-5-20251001 or claude-sonnet-5-5, both currently on sale in /models: with your QCode API key, set the Anthropic-protocol base URL to https://api.qcode.cc/api (the same value goes into ANTHROPIC_BASE_URL for Claude Code) and put the matching ID in the model field. The QCode docs state that Claude models only work through the Anthropic-protocol endpoint, not the OpenAI-compatible one. Usage is billed per token; see /models for each model's rate. If /models later lists claude-haiku-5-5, handle the roughly 30% token difference and the breaking changes from the section above before switching.
Frequently asked questions
When was Claude Haiku 5.5 released?
On 2026-10-07. The official model page reads "Released October 7, 2026", and the Claude API release notes and Anthropic's launch post went up the same day; its status is Active, with retirement no sooner than 2027-10-07.
How much does Claude Haiku 5.5 cost?
It has two tiers by prompt length. For prompts up to 100,000 tokens: input $0.10, output $0.50, cache reads $0.01, 5-minute cache writes $0.125, 1-hour cache writes $0.20 per million tokens. For prompts over 100,000 tokens: input $0.50, output $2.50, cache reads $0.05, 5-minute cache writes $0.625, 1-hour cache writes $1. The Batch API takes a further 50% off input and output.
How much more does a prompt over 100K cost?
Every rate is five times the lower tier, converted from the official price list. With 2,000 output tokens in both cases, a request with 50,000 input tokens costs $0.005 + $0.001 = $0.006, while one with 150,000 input tokens is billed entirely at the higher tier: $0.075 + $0.005 = $0.08.
What is the model ID on each platform?
claude-haiku-5-5 on the Claude API, Google Cloud, Microsoft Foundry and Claude Platform on AWS, and anthropic.claude-haiku-5-5 on Amazon Bedrock. The migration guide says claude-haiku-5-5 is a fixed model ID with no date suffix and no separate alias.
How is Haiku 5.5 different from Haiku 4.5?
Mainly in four ways: price (converted from the official price list, one tenth of Haiku 4.5 per token up to 100K and half above it), limits (1M / 128K vs 200K / 64K), thinking (adaptive thinking on by default with effort, default medium; manual budget_tokens returns 400) and tokenization (about 30% more tokens for the same text). Haiku 4.5's official retirement date is no sooner than 2026-10-15 (a floor, not a shutdown date), and its status is still Active.
Is Haiku 5.5 on QCode?
Check /models: that live list shows whether claude-haiku-5-5 is on QCode. Today you can use claude-haiku-4-5-20251001 or claude-sonnet-5-5 from /models with the Anthropic-protocol base URL https://api.qcode.cc/api, billed per token at the rates shown on /models.
Sources
Specs, per-platform IDs, both price tiers and the retirement commitment: the Claude Haiku 5.5 overview page, What's new in Claude Haiku 5.5 and the migration guide in Anthropic's official docs (crawled 2026-10-08). Full price list, the long-context pricing note and the Batch discount: the official pricing page, crawled the same day. Release date and platforms: the 2026-10-07 entry in the Claude Platform release notes and the Anthropic news page. Positioning, use cases and the footnote behind "around 75% less": the official launch post Introducing Claude Haiku 5.5 and the Claude Haiku product page. Default effort and the limits on turning thinking off: the official effort docs and Prompting Claude Haiku 5.5. Haiku 4.5 specs and retirement date: the official Claude Haiku 4.5 overview page. Sonnet 5.5 specs: the official Claude Sonnet 5.5 overview page. QCode connection address: the QCode docs page on endpoints and API formats (updated 2026-09-25).
Start with the Claude models on sale today
One key for Claude models such as claude-haiku-4-5-20251001 and claude-sonnet-5-5, billed per token; whether Haiku 5.5 is listed is shown on /models.
Related reading
Claude Haiku 5.5: release timing and what is known
The tracker covering the run-up to launch and what was known when.
Claude Haiku 4.5: the complete guide
Specs, pricing and usage of the previous Haiku.
Claude Sonnet 5.5 complete guide
1M context, $2 / $10 and the breaking changes.
The figures and quotes on this page were checked on 2026-10-08 against Anthropic's official docs, price list and launch post; upstream changes may happen without notice. Cost claims such as "around 75% less" are the vendor's own statements and calculations, not a promise about your specific workload. Model availability is whatever /models shows.