
Claude Haiku 5.5 needs a breakpoint cost test
In short: Claude Haiku 5.5's low base rate has a fivefold prompt-length breakpoint; re-tokenise real workloads before routing production traffic.
On 7 October 2026, Anthropic released Claude Haiku 5.5 with a one-million-token context window and a base API price of $0.10 per million input tokens and $0.50 per million output tokens. Those numbers apply only when the prompt is at or below 100,000 tokens.
My finding is narrower than “cheap model, route more traffic”. Claude Haiku 5.5 needs a breakpoint cost test against your real prompt distribution, because crossing 100,000 prompt tokens multiplies every listed rate by five, while its tokenizer can count the same text differently from Haiku 4.5.
Why Claude Haiku 5.5 needs a breakpoint cost test
Anthropic's model overview lists two price bands. The larger context window is technically available in both, but it is not economically flat.
| Claude Haiku 5.5 charge | Prompt at or below 100K | Prompt above 100K | Multiplier |
|---|---|---|---|
| Input, per million tokens | $0.10 | $0.50 | 5x |
| Output, per million tokens | $0.50 | $2.50 | 5x |
| Five-minute cache write | $0.125 | $0.625 | 5x |
| Cache read | $0.01 | $0.05 | 5x |
These are Anthropic's public list rates on 9 October 2026; they are not a measured production bill.
Take a synthetic request with 90,000 input tokens and 10,000 output tokens. Ignoring cache charges, its listed model cost is $0.014. At 110,000 input tokens and the same output, the listed cost becomes $0.080. The request grew by 20,000 input tokens; the bill grew by more than five times because it also changed bands.
That arithmetic is not a benchmark. It is a reason to stop storing “$0.10/$0.50” as the model's single price.
The Haiku 5.5 tokenizer can move the boundary
Anthropic also says Haiku 5.5 uses the newer tokenizer shared by Claude 4.7 and later models. The official documentation warns that the same text counts as approximately 30% more tokens than on Haiku 4.5.
A workload measured at roughly 77,000 Haiku 4.5 tokens could therefore land near 100,000 Haiku 5.5 tokens before anyone adds a tool result, retrieved document or retry context. That is an approximation, not a conversion guarantee. The safe move is to re-tokenise representative requests with the target model's tokenizer and inspect p50, p95 and p99 prompt lengths.
The one-million-token headline still matters. It tells you what the model can accept. It does not tell you which price function a production request will enter.
My routing-table correction
This is editorial judgment, not an account of an AndyMental, Trigent or client deployment. I have not benchmarked Haiku 5.5 quality, latency or realised cache behaviour.
I would replace the scalar price cell in a routing catalogue with a versioned function: model ID, tokenizer version, prompt-length band, input/output rates, cache mode, effort setting and effective date. Then I would replay the workload histogram across that function before comparing it with another route.
This advances my earlier rule that a model-routing price needs an expiry date. Time is one coordinate. Haiku 5.5 makes prompt length and tokenizer identity two more.
Where this holds and where it breaks
The breakpoint test can reject a bad cost forecast. It cannot select the model by itself. Anthropic positions Haiku 5.5 for routing, extraction, classification and subagent work, but those are vendor recommendations, not evidence that it clears your acceptance tests.
Run the cost replay beside task quality, tail latency, retries and human review. A higher token bill can still produce cheaper accepted work. A lower price can still be expensive if the route fails and escalates.
What's in it for you
- Re-tokenise real prompts before carrying Haiku 4.5 token counts into a Haiku 5.5 forecast.
- Report how much traffic sits below, near and above the 100,000-token breakpoint.
- Store price as a versioned function, then pair it with accepted-work results before changing production routes.
A cheap model rate is not one number when the tokenizer and prompt-length band can change which number you pay.
Sources
- Anthropic — Claude Haiku 5.5 model overview, accessed 9 October 2026
- Anthropic — Introducing Claude Haiku 5.5, 7 October 2026


