October 8, 2026. Anthropic released Claude Haiku 5.5 on October 7 at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, 90% below Haiku 4.5's $1 and $5 and the same list price as OpenAI's GPT-6 Luna. Prompts over 100,000 tokens pay $0.50 and $2.50. Anthropic puts the average saving over Haiku 4.5 at about 75% once its new tokenizer and long prompts are counted, per its launch post and pricing page. For agencies running lead routing and summaries at volume, a task now costs a fraction of a cent.

Key numbers
| Item | Number |
|---|---|
| Haiku 5.5 input, prompts up to 100K tokens | $0.10 per million |
| Haiku 5.5 output, prompts up to 100K tokens | $0.50 per million |
| Haiku 5.5 cache reads ($0.05 over 100K tokens) | $0.01 per million |
| Haiku 5.5 input and output, prompts over 100K | $0.50 and $2.50 per million |
| Haiku 5.5 Batch, prompts up to 100K | $0.05 and $0.25 per million |
| Haiku 4.5 input and output | $1 and $5 per million |
| Average saving vs Haiku 4.5, per Anthropic | about 75 percent |
| Tokenizer change vs Haiku 4.5 (for the same text) | about 30 percent more tokens |
| Context window and max output | 1M tokens and 128K tokens |
| GPT-6 Luna input and output, up to 272K tokens | $0.10 and $0.50 per million |
| Sonnet 5.5 cache reads from October 7 (was $0.20) | $0.10 per million |
| 10,000 lead classifications on Haiku 5.5 (1,500 tokens in, 50 out each) | $1.75 |
Prices read on 8 October 2026 from Anthropic's pricing page and Claude Haiku 5.5 launch post and OpenAI's GPT-6 Luna model page.
What Claude Haiku 5.5 costs
- Prompts up to 100,000 tokens: $0.10 input, $0.50 output, $0.125 for 5-minute cache writes, $0.20 for 1-hour writes and $0.01 for cache reads, all per million tokens.
- Prompts over 100,000 tokens: $0.50 input, $2.50 output, $0.625 and $1 for cache writes and $0.05 for cache reads.
- Batch: half price, $0.05 and $0.25 for short prompts and $0.25 and $1.25 for long ones.
- Against Haiku 4.5: $1 input, $5 output and $0.10 cache reads, so short prompts are 90% cheaper and long ones 50% cheaper.
- Sonnet 5.5 got cheaper the same day: its cache reads fell from $0.20 to $0.10 per million, which Anthropic says makes it about 20% cheaper on most agentic work.
VentureBeat reported the same prices and the match with GPT-6 Luna. The model has a 1 million token context window and up to 128,000 output tokens. It runs on the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry under the ID claude-haiku-5-5.

The 100,000-token line and the tokenizer
Two details decide whether a migration saves 90% or much less. First, Haiku 5.5 is priced by prompt length: a request whose prompt runs over 100,000 tokens pays the higher input, output and cache prices. No other Claude model from 4.6 on is priced this way. Second, Haiku 5.5 uses the newer tokenizer of Claude 4.7 and later, so the same text counts as about 30% more tokens than on Haiku 4.5. A document that measured 80,000 tokens on Haiku 4.5 comes to roughly 104,000 on Haiku 5.5 and crosses the line. Anthropic says about 90% of requests to Haiku 4.5 were under 100,000 tokens, but recount your largest prompts before you switch.
Code needs changes too: per Anthropic's migration notes, temperature, top_p, top_k and assistant prefill now return errors, manual extended thinking with budget_tokens returns a 400 error, and adaptive thinking is on by default.
Claude Haiku 5.5 vs GPT-6 Luna on price
On short prompts the list prices match exactly: $0.10 input, $0.50 output, $0.01 cache reads and $0.125 cache writes on both. They split on long prompts. OpenAI's GPT-6 Luna steps up only past 272,000 input tokens, to 2x input and 1.5x output for the full request, or $0.20 and $0.75. Haiku 5.5 steps up past 100,000, to 5x. A 150,000-token document with a 1,000-token answer costs about $0.016 on Luna against $0.078 on Haiku 5.5 at list prices, before tokenizer differences. Anthropic's own benchmarks put Haiku 5.5 ahead of Luna on knowledge work and coding, and HubSpot told Anthropic that Haiku 5.5 scored 92.8% on its CRM task suite, the best result HubSpot had seen on it.
What it means for operators
At list price, a lead-routing call with 1,500 input tokens and a 50-token answer costs about $0.000175 on Haiku 5.5, so tagging 10,000 inbound leads costs $1.75. A support or call summary with 3,000 tokens in and 300 out costs $0.00045, or $4.50 per 10,000. Tagging every lead by service and urgency, scoring every cold email reply or summarizing every call is now cheap enough to run on every record. Intake is the obvious fit: a law firm that sorts every web and phone enquiry by practice area and urgency, the triage our AI intake for law firms runs, can do it on a small model and keep a larger one for drafting. Test on a few hundred of your own records first; Anthropic still recommends Sonnet 5.5 and Opus 5.5 for complex agentic coding. If you pay for Claude Max or Team, the new monthly API credit covers this kind of workload, and our AI automation team can wire it into your CRM.
Frequently Asked Questions
For prompts up to 100,000 tokens: $0.10 per million input tokens, $0.50 per million output tokens, $0.125 for 5-minute cache writes, $0.20 for 1-hour writes and $0.01 for cache reads. Prompts over 100,000 tokens pay $0.50 input and $2.50 output. Batch processing is half price.
On prompts up to 100,000 tokens the list prices are the same: $0.10 input and $0.50 output per million tokens. Between 100,000 and 272,000 input tokens GPT-6 Luna is cheaper, because Haiku 5.5 moves to $0.50 and $2.50 while Luna stays at its standard rate until 272,000. The tokenizers also count text differently, so test on your own data.
Two reasons. A prompt over 100,000 tokens pays five times the short-prompt price, and Haiku 5.5's newer tokenizer counts the same text as about 30% more tokens than Haiku 4.5, so prompts that sat just under 100,000 tokens before can cross the line. Recount your largest prompts before migrating.
Per Anthropic's migration notes, setting temperature, top_p or top_k returns an error, assistant message prefill returns an error, and manual extended thinking with budget_tokens returns a 400 error. Adaptive thinking is on by default, and responses can also begin with thinking blocks, so select content blocks by type.