Does the 100K tier apply to the whole prompt or only the portion above 100K?
The coverage I found describes it as prompts above 100K tokens costing $0.50 input and $2.50 output, which reads as the whole request moving to the higher tier rather than a surcharge on the excess only. That is my reading of the reporting's wording, not a verbatim statement from the official pricing page.
Before going live, treat the official pricing page as authoritative and verify with one real 110K-Token request against your billing detail.
Should I switch all existing Haiku 4.5 traffic to Haiku 5.5?
Not necessarily. Short-context, fixed-format classification and summarization is usually worth switching, but if your requests often sit between 77K and 270K tokens, the new tokenizer plus the $0.50 tier shrinks the saving and can even lose to other models' list prices.
The safer approach is to sample a week of real requests, estimate tokens at 1.3 times, price them in the two tiers, and then decide which routes to move and which to keep.
What happens if my existing code sends a custom Temperature after switching models?
Per the coverage, Haiku 5.5 returns a 400 error when it receives a non-default temperature, top_p or top_k, meaning the request fails outright rather than the value being ignored.
If your service has retry or fallback logic, make sure it does not treat this 400 as a transient error and retry forever; the cleanest fix is to remove these parameters before the switch.
Haiku 5.5 beats GPT-6 Luna on the benchmarks, so can I use that to choose a model?
I would not. These figures are published by Anthropic and relayed by the press, the task sets and settings are chosen by the vendor, and Sonnet 5.5 scores higher on every row.
The most reliable basis is a small side-by-side test on your own tasks and prompts, comparing quality together with the real cost per thousand requests rather than leaderboard percentages alone.
Anthropic launched Claude Haiku 5.5 (model ID claude-haiku-5-5) on October 7, 2026, describing it as its cheapest and fastest small model to date, aimed at summaries, context Compaction in Claude Code, classification and Subagent tasks. Claude Code v2.1.293 made it the default Haiku model on the Anthropic API the same week.
Haiku 5.5 offers a 1M-Token Context Window and up to 128K output tokens (batch jobs reach 300K in beta), takes text and images as input, and has a June 2026 Knowledge Cutoff. Pricing has two tiers per million tokens: prompts up to 100K tokens cost $0.10 input and $0.50 output; above 100K it is $0.50 input and $2.50 output. Cache reads cost $0.01, 5-minute cache writes $0.125, and Batch Processing takes 50% off. For comparison, Haiku 4.5 charged $1 and $5. It is also the first Haiku-class model with an adjustable effort setting, defaulting to medium, with Adaptive Thinking on by default; passing a non-default Temperature, top_p or top_k returns a 400 error.
Per press coverage of Anthropic's figures, Haiku 5.5 is 90% cheaper than Haiku 4.5 at short context. Anthropic says about 90% of Haiku 4.5 requests stayed under 100K tokens, and since the new tokenizer counts roughly 30% more tokens for the same text, it estimates about 75% cheaper on average after adjusting. One detail is easy to miss: a prompt of 77K tokens on the old model becomes roughly 100K on the new tokenizer, right at the line. Take a prompt of 100K tokens on Haiku 4.5, which costs $0.10. On Haiku 5.5 it becomes about 130K tokens, billed at the $0.50 tier (inferred from the reported tier description), about $0.065, so the saving shrinks from 90% to about 35%. Conversely, a 50K-token prompt on Haiku 4.5 becomes about 65K, and its cost falls from $0.05 to about $0.0065, a saving of roughly 87%.
Anthropic's reported numbers show OSWorld 2.1 (offline subset) at 72.4% for Haiku 5.5, 48.9% for GPT-6 Luna and 15.7% for Haiku 4.5; Terminal-Bench 4.0 at 39.2% against Luna's 16.4% and Haiku 4.5's 0.0%. Humanity's Last Exam is 45.9% without tools and 57.4% with tools. These are vendor-reported figures, and Sonnet 5.5 leads on every row (70.6% on Terminal-Bench 4.0). Anthropic itself recommends Sonnet 5.5 or Opus 5.5 for complex agentic coding, positioning Haiku 5.5 as a subagent beneath them. On price, GPT-6 Luna has the same short-context list price of $0.10 and $0.50, but its higher tier starts only above 272K input tokens ($0.20 and $0.75), so at around 150K tokens Luna's list price is lower than Haiku 5.5's. Haiku 5.5 is generally available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS.
Before migrating, run two checks on your own traffic. First, multiply your current request token counts by 1.3 and see what share crosses 100K. Second, look for any code that sends a custom temperature, which will now become a 400 error. If your workload is mostly short-context classification, summarization or subagent calls, the savings are substantial; if many requests sit between 100K and 270K, compare Haiku 5.5's $0.50 tier against other models' list prices one by one instead of relying on the $0.10 headline.