Your privacy choices

Optional Google Analytics cookies help us understand which guides are useful. They stay off unless you accept. You can change your choice in the footer. Privacy policy

Claude Haiku 5.5 Cuts Short-Prompt Token Prices by 90%—Here’s the Catch

Haiku 5.5’s lowest API rates are a tenth of Haiku 4.5’s. Longer prompts and a new tokenizer complicate the savings. We work through the numbers.

By Ubedulla · 5 min read
A copper microchip on a branching track leading to different-sized stacks of coins.
AI-generated editorial illustration of AI pricing tiers.

A million input tokens used to cost $1 on Claude Haiku 4.5. For shorter prompts, Haiku 5.5 brings that price down to ten cents.

That is an attention-grabbing cut. It also leaves out the part developers need before estimating their next bill: the new model has two pricing tiers, and the same text can now count as more tokens.

Anthropic released Claude Haiku 5.5 on October 7, targeting frequent, narrowly scoped work such as summaries and classification. Its announcement estimates average running costs around 75% below Haiku 4.5. The cheaper headline rates and that smaller average saving describe different things.

The price cut depends on prompt length

Here are the standard input and output rates in Anthropic’s Haiku product information, in US dollars per million tokens:

Model and prompt tierInputOutput
Haiku 4.5$1.00$5.00
Haiku 5.5: up to 100,000 prompt tokens$0.10$0.50
Haiku 5.5: above 100,000 prompt tokens$0.50$2.50

The lower tier is 90% cheaper per token than Haiku 4.5. The upper tier is 50% cheaper. Neither percentage promises an identical reduction in the cost of completing a real job.

For a developer sorting incoming support tickets, this is worth investigating. For someone repeatedly sending an entire company handbook alongside each ticket, the prompt-length boundary deserves equal attention.

What would 100,000 small requests cost?

Consider an illustrative workload: 100,000 requests, each consuming 2,000 input tokens and producing 200 billable output tokens. Assume those counts are identical on both models, with no caching, batch discounts, tool fees, or retries.

That adds up to 200 million input tokens and 20 million output tokens. Our calculation is:

  • Haiku 4.5: $200 for input plus $100 for output, or $300.
  • Haiku 5.5’s lower tier: $20 for input plus $10 for output, or $30.

The $270 difference is real arithmetic, not a measured production saving. The assumptions do most of the work. Change how much the model writes, how often it needs another attempt, or which tier a request reaches, and the total changes.

That is why a useful trial keeps two columns: dollars spent and tasks completed correctly. A cheap answer that sends a customer to the wrong team still creates work for somebody.

The catch: your old token counts are no longer a safe budget

Anthropic’s migration guide says the new tokenizer produces approximately 30% more tokens for the same input text, with the increase depending on the content. It tells developers to count prompts again using the new model.

Suppose a prompt previously measured 80,000 tokens. A hypothetical 30% increase would put it at 104,000—past the new pricing boundary. That is an illustration of the risk, not a prediction for every 80,000-token document.

There are integration changes too. The guide replaces manually budgeted thinking with adaptive thinking and warns about previously accepted request parameters. A migration can therefore require more than changing the model name.

My suggested first step would be to replay a small, representative batch in a test environment. Include the longest prompts you actually use. Record the new input counts, all billable output, failures and elapsed time. Only then extrapolate a monthly bill.

A smaller model gets a choice about how hard to think

Haiku 5.5 is the first Haiku release with adjustable effort. Anthropic positions that control as a way to trade cost against capability, and still recommends its larger models for difficult agentic coding work. The launch page makes that distinction explicitly.

For an app builder, the interesting experiment is to separate jobs. Extracting an order number from a short message and investigating a broken checkout system are different tasks. They do not automatically deserve the same model or effort setting.

Imagine a support tool that asks a small model to identify the product, topic and urgency, then routes ambiguous cases to a person. The useful measure is whether the route was right. Adding a paragraph of confident explanation does not make that decision more correct.

Before giving such a tool permissions, test the cases that look deceptively easy: two order numbers in one email, an angry message about an already-resolved issue, or instructions copied from somebody else’s conversation. These are proposed evaluation cases, not tests we have run on Haiku.

Who should pay attention?

Teams with lots of small AI calls have the clearest reason to investigate. They can compare an existing workload against a new candidate without rebuilding the whole product around launch-day claims.

For people using Claude mainly through its chat app, API token prices are a separate issue from subscription charges. A lower API rate does not itself establish that your monthly chat plan has become cheaper. Our AI glossary explains terms such as tokens if you are encountering these pricing tables for the first time.

The promising part of this release is what a lower cost per successful task might make practical. The question to answer is specific: which repeated job could you now afford to improve, and how would you know the improvement worked?

Quick answers

Is Haiku 5.5 always 90% cheaper?

No. That comparison applies to the listed per-token rates in the lower prompt tier. Actual costs also depend on token counts, output, retries and other charges.

Have we tested it?

No. This is a source-based pricing analysis with illustrative calculations, not a hands-on benchmark.

Researched October 8, 2026. Sources: Anthropic’s announcement, Haiku product page and migration guide. Calculations and suggested test cases are The Bot Post’s analysis. AI-assisted, source-checked analysis. Prices checked October 8, 2026.

About the author

Ubedulla

Founder & Editor

Founder and editor of The Bot Post, covering AI news and technology.

Related Articles