The short version: Claude Haiku 5.5 pricing starts at $0.10 per million input tokens and $0.50 per million output tokens, a tenth of what Claude Haiku 4.5 charges, according to Anthropic’s price list (checked October 7, 2026, the day the model launched). Those prices hold for any prompt up to 100,000 tokens, which is about 55,000 words. Past that line, the whole request, the reply included, bills at $0.50 and $2.50: five times as much. On either side it is still the cheapest Claude by a wide margin, and it now handles office work its predecessor could barely touch. The one habit worth building is keeping what you send it under the line.
That line is new for Claude. Every other current Claude model charges one flat rate whatever the prompt length; Anthropic’s pricing page names Haiku 5.5 as the single exception. So the cheapest model in the lineup is also the only one where the size of what you feed it changes the price per token.
What does Claude Haiku 5.5 cost?
Anthropic released Claude Haiku 5.5 on October 7, 2026, as its fastest and smallest model, aimed at high-volume work like sorting, extracting and routing. The prices, read off the pricing page the same day:
Prompts up to 100,000 tokens: $0.10 per million input tokens, $0.50 per million output tokens, $0.01 per million for cached input.
Prompts over 100,000 tokens: $0.50 input, $2.50 output, $0.05 cached.
Haiku 4.5, for comparison: $1 input, $5 output, at any length.
Anthropic’s headline says Haiku 5.5 costs “around 75% less to run” than Haiku 4.5. One adjustment sits in the documentation instead. Haiku 5.5 uses the newer tokenizer that Claude models adopted with version 4.7, and Anthropic’s what’s new page says the same text counts as about 30 percent more tokens than on Haiku 4.5. So for the same email or the same document, a short request costs 1.3 times $0.10, or $0.13, for every dollar it used to cost on Haiku 4.5. That is 87 percent cheaper rather than 90. Above the line it is 1.3 times $0.50, or $0.65 per dollar: 35 percent cheaper rather than 50. Still a cut at every length, just not the round number.
The good news: the cheap model can now do real office work
The more interesting change is not the price. It is what the price now buys. On Anthropic’s own published benchmarks, Haiku 5.5 scored 72.4 percent on OSWorld 2.1, a test of operating a real computer through screens, menus and forms. Haiku 4.5 scored 15.7 percent. Sonnet 5.5, a model that costs twenty times as much per input token, scored 83.9. On chart reading, Haiku went from 6.4 percent to 46.4. On GDPval-AA, a set of tasks drawn from real occupations, it went from 735 to 1,620, against Sonnet 5.5’s 1,840.
These are the vendor’s numbers, and the only test that counts is your own work. But the direction is plain. The chores an office hands to a fast, cheap model, like reading every incoming email and tagging it as a quote request, a complaint or a scheduling change, pulling the job address and phone number out of a message, or filling in a supplier’s web form, used to sit at the edge of what the smallest Claude could do reliably. They now sit well inside it. That is time handed back to whoever currently does that sorting by hand, at a cost measured in pennies a month.
It is the same shift we saw with Claude Sonnet 5.5 on the free plan, one rung lower: work that needed the middle model months ago is moving to the bottom one.
Where is the 100,000-token line, in pages?
Anthropic’s models page says 1 million tokens is roughly 555,000 words on the current tokenizer. That puts the 100,000-token line at about 55,500 words, or around 110 single-spaced pages at 500 words a page.
A single customer email plus a page of instructions is a few thousand tokens. You will never get near the line that way. Three things get you there:
Attaching everything to every request: the full price book or handbook, sent with each question so the assistant “knows the business.”
Long conversations: each new turn in a chat or agent session resends the history before it, so a thread that started small grows until it crosses the line.
Computer use: Anthropic’s pricing page says the computer use toolset adds about 4,500 input tokens to each request, before screenshots, which bill as images on top.
The price list also tiers cache reads and cache writes by the same prompt length, which tells you that text you have cached still counts toward the size of the prompt. Caching makes the tokens cheaper. It does not make the prompt shorter.
What does crossing the line do to a bill?
Start with one request. A 99,000-token prompt with a 1,000-token reply costs $0.0099 for the input and $0.0005 for the reply: about one cent. Make the prompt 101,000 tokens, two pages longer, and the input costs $0.0505 and the reply $0.0025. That is 5.3 cents. Cross the line and every token in the request, the reply included, is billed at five times the rate.
Now a month. Picture a plumbing company with three vans and an office assistant built on Haiku 5.5 that answers questions from the techs and the front desk: does this job need a permit, what do we charge for a water heater swap, what is the warranty on a repipe. Thirty questions a working day, 22 working days, so 660 questions. Each answer runs about 400 tokens, counting any thinking the model does, since thinking bills as output.
The whole manual attached to every question, 120,000 tokens: 120,500 input tokens at $0.50 per million plus 400 output tokens at $2.50 per million is about 6.1 cents a question, or about $40 a month.
The same manual trimmed to 95,000 tokens: 95,500 input tokens at $0.10 plus 400 output at $0.50 is just under 1 cent a question, or about $6.40 a month.
The original 120,000-token setup on Sonnet 5.5: about 24.5 cents a question, or about $162 a month.
Read together, the line looks less like a trap than a discount for tidiness. On the wrong side of it, Haiku 5.5 still costs a quarter of Sonnet. On the right side, it costs a twenty-fifth. The difference for this plumber is about $34 a month, which is real money for a few minutes spent deciding what the assistant actually needs to read.
We ran the same kind of real-month arithmetic on Grok 4.7’s thinking tokens and on DeepSeek’s off-peak hours. Both times a setting most owners never see decided the bill. This is the third.
Has anyone priced a model this way before?
Google has. Its Gemini API price list charges more for Gemini 2.5 Pro and Gemini 3.1 Pro Preview once a prompt passes 200,000 tokens. On 3.1 Pro Preview, input goes from $2 to $4 per million and output from $12 to $18. Haiku 5.5’s line sits half as far out, and the step is steeper: five times on both input and output, where Gemini doubles input and adds half again to output. The idea is familiar; the size of the jump is not.
Three habits that keep you on the cheap side
Send the chapter, not the book. If an assistant answers questions from a long document, have whoever set it up pull only the relevant section into each request. Many tools that “chat with your documents” are built to do exactly this; it is worth asking whether yours is.
Start a fresh conversation for a new job. One endless thread is the easiest way to cross the line without noticing. For automated agents, ask your builder to cap or summarize the history rather than resend all of it.
Ask which side of the line your requests land on. Every response from Anthropic’s API reports the number of input tokens it used, so a developer or the automation tool you pay for can tell you in one look. And if you switch an existing automation to Haiku 5.5 and a step suddenly fails, check its temperature setting first: Anthropic’s migration notes say Haiku 5.5 returns an error for any non-default temperature, top_p or top_k value.
And for anything that can wait until morning, like tagging yesterday’s emails, Anthropic’s Batch API takes another 50 percent off.
Frequently Asked Questions
How much does Claude Haiku 5.5 cost?
For prompts up to 100,000 tokens, Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens, per Anthropic’s pricing page on October 7, 2026. For prompts over 100,000 tokens it costs $0.50 and $2.50. The Batch API takes 50 percent off both.
What happens when a Claude Haiku 5.5 prompt goes over 100,000 tokens?
The whole request moves to the higher price tier, so input, output and cache prices all rise fivefold for that request. A 101,000-token prompt with a 1,000-token reply costs about 5.3 cents, against about one cent for a 99,000-token prompt with the same reply.
Is Claude Haiku 5.5 cheaper than Haiku 4.5 for the same text?
Yes, at every length. Because Haiku 5.5 counts the same text as about 30 percent more tokens, the real saving on the same text is about 87 percent for short prompts and about 35 percent for prompts over the 100,000-token line, rather than the 90 and 50 percent the price per token suggests.
When should a small business use Claude Sonnet 5.5 instead?
Use Haiku 5.5 for short, repeated jobs like sorting, tagging, extracting details and simple form filling, and try Sonnet 5.5 when the work needs more judgment, longer reasoning or harder coding. On Anthropic’s own benchmarks Sonnet 5.5 still leads, for example 83.9 percent to 72.4 on computer use, at twenty times the input price.
Where in your business does the work arrive in small, repeatable pieces, and where does it arrive as one enormous file? Tell us in the comments.
