Claude Haiku 5.5 Is Cheap Until Long Prompts Trigger a 5x Rate

Claude Haiku 5.5 starts at $0.10 per million input tokens, but prompts over 100K cost five times more. Here is how to budget and migrate.
Claude AI symbol used for coverage of Anthropic models, APIs, and agent tools.
Claude AI symbol. Image source: Wikimedia Commons; public-domain/CC0 asset.

Anthropic released Claude Haiku 5.5 on October 7, cutting the list price of its small model by as much as 90% and giving it adaptive reasoning, a one-million-token context window and stronger computer-use performance. It is available now through Anthropic’s API, Amazon Bedrock, Google Cloud and Microsoft Foundry.

The headline price is compelling: $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. But production teams should pay close attention to the boundary in that sentence. Once a prompt exceeds 100,000 tokens, both rates jump fivefold. A tokenizer change also means the same text can count as roughly 30% more tokens than it did with Haiku 4.5, according to Anthropic’s model documentation.

Claude AI symbol used for Anthropic model and API coverage.
Claude Haiku 5.5 is built for high-volume, latency-sensitive API work. Image: Wikimedia Commons (public domain/CC0)

That combination makes Haiku 5.5 unusually inexpensive for short, repetitive work, while creating a sharp cost cliff for long-context agents and document pipelines. Developers evaluating the model should test their real prompt-size distribution, not estimate savings from the lowest advertised rate alone.

Claude Haiku 5.5 pricing changes at 100,000 tokens

Anthropic uses two pricing tiers based on the length of the input prompt. The higher tier applies to requests whose prompts exceed 100,000 tokens.

Usage Prompt up to 100K Prompt over 100K
Input $0.10 per million tokens $0.50 per million tokens
Output $0.50 per million tokens $2.50 per million tokens
5-minute cache write $0.125 per million tokens $0.625 per million tokens
1-hour cache write $0.20 per million tokens $1.00 per million tokens
Cache read $0.01 per million tokens $0.05 per million tokens

The threshold is a request-level distinction, not a gradual surcharge on only the tokens beyond 100,000. At list price, a request with a 100,000-token prompt and 5,000-token response costs about 1.25 cents. A prompt one token longer, with the same response size, costs about 6.25 cents. Batch processing cuts input and output charges by 50%, but it does not remove the two-tier structure.

Haiku 4.5 cost $1 per million input tokens and $5 per million output tokens. That makes Haiku 5.5’s short-prompt rates 90% lower on paper. Anthropic estimates a more modest average reduction of about 75% per task because the newer model’s tokenizer consumes more tokens for the same text and because some traffic falls into the higher tier. The company says about 90% of requests to the previous Haiku stayed below 100,000 tokens.

The tokenizer can move workloads across the price line

Haiku 5.5 uses the newer tokenizer found in recent Claude models. Anthropic’s documentation warns that identical text counts as approximately 30% more tokens than on Haiku 4.5. A workload that previously occupied 80,000 tokens could therefore land near 104,000 before any other prompt changes, although the precise count depends on the content.

This matters most for systems that assemble prompts dynamically. Retrieval-augmented generation pipelines may add search results, policy text, tool schemas and conversation history before a user question reaches the model. Coding agents may include repository maps, error logs and file contents. Browser agents can accumulate screenshots, accessibility trees and action history. A modest increase in any one component can push the complete request into the higher tier.

Teams migrating from Haiku 4.5 should log input tokens per request and examine the 90th, 95th and 99th percentiles. Averages can hide a small share of long prompts that produce a disproportionate part of the bill. Useful controls include capping retrieved chunks, summarizing old conversation turns, pruning unused tool definitions and routing long-context tasks to a model selected for that workload rather than sending everything through one endpoint.

Where Haiku 5.5 is designed to fit

Anthropic positions Haiku 5.5 for classification, extraction, routing, summarization, database queries, live support and narrowly defined subagent work. The model accepts text and images, supports a one-million-token context window, and can return up to 128,000 output tokens in normal API use. A beta header raises the maximum output to 300,000 tokens for the Message Batches API.

It is also the first Haiku model with adaptive thinking. Developers can adjust an effort parameter instead of choosing one fixed reasoning depth for every request. Medium is the default. That gives teams another routing lever: low-effort execution for predictable extraction or labeling, and more reasoning for tasks that justify extra latency and tokens.

Migration is not purely a model-name swap. The API model ID is claude-haiku-5-5. Anthropic advises developers to omit non-default values for temperature, top_p and top_k; supplying one returns an HTTP 400 error. Thinking blocks are also scoped to the account that generated them, or to a linked account, which matters for systems that persist and replay model state across environments.

On Amazon Bedrock, the global inference-profile model ID is global.anthropic.claude-haiku-5-5. AWS also offers US, EU, Australia and Japan geographic inference profiles, plus GovCloud availability. Bedrock customers can use familiar IAM, CloudTrail, CloudWatch and Guardrails controls, while Anthropic’s native platform is separately available through the AWS console.

Benchmarks show a large jump, with important limits

Anthropic reports 72.4% on the offline subset of OSWorld 2.1, a computer-use benchmark, up from 15.7% for Haiku 4.5. On Terminal-Bench 4.0, which measures multi-step command-line work, Haiku 5.5 scored 39.2%, compared with 70.6% for Sonnet 5.5. Its reported GDPval-AA v2.1 knowledge-work score was 1,620, between GPT-6 Luna at 1,437 and Sonnet 5.5 at 1,840.

Those are vendor-published results and should not substitute for workload-specific evaluation. Anthropic itself recommends Sonnet 5.5 and Opus 5.5 for complex agentic coding, while framing Haiku as the execution layer for narrower assignments. Early customers reported gains on CRM audits, document questions and enterprise summaries, but those tests were selected for the launch and used different internal scoring methods.

The practical comparison is therefore not simply whether Haiku beats a larger model on one benchmark. It is whether a cheaper model can clear the acceptance threshold for a tightly specified task, including error rate, latency, tool-call reliability and the cost of retries or human review.

What teams should test before switching

  • Measure token counts again. Re-run representative production prompts with the new tokenizer and chart how many cross 100,000 tokens.
  • Separate short and long routes. Classification and extraction traffic may suit Haiku, while ambiguous planning and large coding tasks remain better candidates for Sonnet or Opus.
  • Evaluate by effort level. Compare quality, latency and total tokens at low, medium and high effort instead of assuming the default is optimal.
  • Test tool behavior. Include malformed responses, timeouts, permission failures and recovery paths, not just successful demonstrations.
  • Price cache policy explicitly. Reused context is cheap, but cache writes and reads also increase fivefold for prompts over the threshold.
  • Check safety fit. Haiku 5.5 allows a wider range of defensive cybersecurity work than Sonnet 5.5’s general safeguards, but still blocks penetration testing and other techniques Anthropic considers more likely to enable abuse.

Haiku 5.5 gives developers a credible new option for high-volume AI features and agent subroutines. Its best economics belong to workloads that are short, repeatable and well measured. For long-context systems, the headline rate is only the starting point: prompt distribution, tokenizer behavior, cache design and routing discipline determine the real bill.

Sources

Leave a Reply

Your email address will not be published. Required fields are marked *

Previous Post
Ethernet cables connected to network switches in a server room

FortiBleed Lockouts Hit FortiGate Firewalls: What Admins Should Check

Next Post
Google SynthID Detector checking a video for an invisible AI watermark

Google’s SynthID Detector Is Public. Here’s What “Not Detected” Means

Related Posts