Claude Haiku 5.5: Pricing and When to Switch

Claude Haiku 5.5 pricing, the 100K prompt threshold, and a monthly bill versus Haiku 4.5 and Sonnet 5.5. Which work should you move?

Published

Claude Haiku 5.5: Pricing and When to Switch

Claude Haiku 5.5, Anthropic's small model for frequent automated tasks, costs $60 a month for one million short classification calls if each uses 500 input tokens and 20 billed output tokens. It is the cheapest current Claude, with rates that rise once a prompt exceeds 100,000 tokens. Evaluate it for classification, extraction, routing and summaries; keep Claude Sonnet 5.5, Anthropic's larger general-purpose model, for work where errors or complex coding outweigh the token saving.

Verified on 8 October 2026 against Anthropic's announcement, pricing documentation and models overview. The budgets below are calculated examples with stated assumptions. All prices are Anthropic's USD list prices.

Claude Haiku 5.5: What Changed

Haiku 5.5 makes narrowly defined AI work much cheaper to run repeatedly. Anthropic released it on 7 October 2026 with the API ID claude-haiku-5-5 and describes it as its cheapest small model yet. Haiku is the small-model line in the Claude family: the starting candidate when you need frequent automated decisions and can define an acceptable answer precisely. Release announcement.

The change from Claude Haiku 4.5, the previous small model, combines lower rates with more capacity and adjustable reasoning. That matters for a founder paying for every ticket label or extracted field. It matters less for a product whose expensive part is correcting the model's mistakes.

Anthropic lists availability through the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry. Its announcement describes the Microsoft availability as Azure. Use your platform's integration and model identifier when moving an existing service. Official platform list.

Claude Haiku Context Window

Haiku 5.5 has a 1M-token context window and 128K maximum output tokens in the standard API. The context window is the material the model can hold for a request; maximum output caps what it can generate. Haiku 4.5 offered 200K context and 64K output. The larger limits accommodate longer work, but they are capacity limits, not a token allowance or a quality guarantee. Current specifications, version changes.

This is also the first Haiku with an adjustable effort setting. Adaptive thinking lets the model decide when and how much to reason, and is enabled by default. Haiku 5.5's API default effort is medium. Announcement, default and thinking mode.

For a simple classifier, evaluate low effort first. For extraction with exceptions or stricter instructions, compare medium and high. Anthropic supports low, medium, high, xhigh and max; it recommends using the highest settings where evaluation shows a gain and comparing them with Sonnet. More reasoning can consume more tokens and time, so record effort alongside the model name in your budget. Effort guidance.

Claude Haiku Pricing: The 100K Prompt Threshold

A prompt of 100,000 tokens gets the low rate; a longer prompt gets the high rate for that request. The threshold concerns the length of each prompt, regardless of how many tokens your account uses during the month.

Tokens are the pieces of text a model processes. Input is what you send, including relevant instructions, history and tool material; output is what the model generates. MTok means one million tokens. These are per-MTok rates:

Model and prompt lengthFresh inputOutputCache hit
Haiku 5.5, up to 100,000 tokens$0.10$0.50$0.01
Haiku 5.5, over 100,000 tokens$0.50$2.50$0.05
Haiku 4.5$1.00$5.00$0.10
Sonnet 5.5$2.00$10.00$0.10

Source: Anthropic's model pricing. A cache hit is a reused part of the input; its price does not apply to every token in a request.

Anthropic's official API pricing documentation showing model token rates and prompt caching prices
Claude API pricing from Anthropic's official documentation

The Haiku boundary is a price step, rather than a higher charge only on the extra tokens. For an illustrative uncached request with 1,000 output tokens, our calculation from the published tiers is:

  • 100,000 input tokens: 0.1 × $0.10 + 0.001 × $0.50 = $0.01050.
  • 100,001 input tokens: 0.100001 × $0.50 + 0.001 × $2.50 = $0.05250050.

One additional input token moves the request onto rates five times higher. Both the input and the output rates change. This is a calculation from the two-tier rate card, with uncached input and the same output assumption.

A clay library shelf steps upward between prompts of 100,000 and 100,001 tokens, with input and output rates rising from $0.10/$0.50 to $0.50/$2.50 per million tokens
Crossing 100,000 prompt tokens changes the request's input and output rates.

For a summarization service, keep the relevant document and instructions together, then count that complete request. Appending an entire conversation or unrelated document collection can change the economics. Shorten or retrieve context when it preserves the answer's quality. Splitting a document blindly can lose the relationship you needed the model to explain.

Batch Cuts Both Token Rates in Half

Use batch when the result can arrive asynchronously, such as overnight tagging or a backlog of document summaries. Haiku 5.5 batch rates are $0.05 input/$0.25 output per MTok for prompts up to 100,000 tokens, and $0.25/$1.25 above it. The Batch API discounts input and output by 50%. Batch rate card, batch processing.

Keep live request routing on a response path that meets your application's deadline. Batch is an adoption choice for work that can wait; it changes delivery as well as price.

A $0.01 Cache Hit Needs an Eligible Reused Prefix

Caching stores a repeated prompt prefix, such as an unchanged extraction policy, for reuse across requests. Haiku 5.5's low-tier read is $0.01 per MTok, after a paid write: $0.125 per MTok for a five-minute cache, or $0.20 for one hour. Above the prompt-length threshold, those prices are $0.05 reads, $0.625 five-minute writes and $1 one-hour writes. Caching prices.

The practical floor is 512 tokens in the cacheable prefix for Haiku 5.5. Shorter prefixes run without caching. A changing customer message is still fresh input, and a newly generated answer is still output. Check cache_creation_input_tokens and cache_read_input_tokens to confirm what happened. Cache limitations.

That distinction matters for short classifiers. The entire 500-token prompt used in the next example is below the caching minimum. A budget that bills those calls as cache hits would understate spend.

Claude Haiku Cost: One Million Monthly Classifications

The example monthly bill is $60 on Haiku 5.5, $600 on Haiku 4.5 or $1,200 on Sonnet 5.5. This is a rate comparison for a SaaS support-ticket classifier that returns a queue label and a short structured decision.

Assume one million calls per month, each consuming 500 fresh input tokens and 20 billed output tokens. Assume these billed counts separately for each model. Every prompt stays below the price threshold. There is no caching, no separately charged server tool and no retry in this baseline.

Monthly consumption is 500 million input tokens and 20 million output tokens. The arithmetic uses the official list rates:

ModelInput chargeOutput chargeMonthly total
Haiku 5.5500 × $0.10 = $5020 × $0.50 = $10$60
Haiku 4.5500 × $1 = $50020 × $5 = $100$600
Sonnet 5.5500 × $2 = $1,00020 × $10 = $200$1,200

Haiku 5.5 costs $0.06 per 1,000 classifications under these assumptions. The same job through batch costs $30, $300 or $600, respectively. Those totals cover model tokens; your application infrastructure and any human review remain separate.

Claude Haiku 4.5 Pricing

Haiku 4.5's $1 input/$5 output per MTok is ten times the new short-prompt rate. That produces the $540 monthly saving in this example. For Haiku 5.5 prompts above 100,000 tokens, the rate difference narrows to half of Haiku 4.5's rates. Published prices.

Keep the equal-token assumption visible. Anthropic says Haiku 5.5's updated tokenizer uses slightly more tokens per task. A tokenizer is the system that divides text into countable tokens. Recount the same input using the new model and measure its billed output, including thinking, before forecasting your own job's saving. Anthropic's tokenizer note, token counting.

Twenty output tokens is an assumption, not a limit to copy into every request. Thinking consumes output capacity too. If the model needs reasoning before returning the label, actual billed output can exceed the short visible answer. Evaluate a lower effort setting and an output cap that leaves enough room to finish. Thinking and output limits.

For other vendors, use the same workload and acceptance criteria with the cheapest AI API comparison. A lower advertised rate is useful when it produces an accepted result within your budget.

What This Means for Builders, Operators and Buyers

Move the repeatable part of the job first. Classification, extraction, routing, summaries and subagents are useful candidates because you can narrow their inputs and inspect their outputs. Anthropic positions Haiku for these kinds of high-volume work. Model positioning.

Builders: Give Haiku a Small Contract

For a funded founder, start with the support-ticket classifier from the bill above. Define the allowed queues, include the ticket, and require one decision. Provide a destination for an unresolved case. A validation check can reject a label outside the allowed set; a scored sample can reveal valid-looking labels sent to the wrong queue.

A solo developer can use Haiku for a similarly narrow extraction task: pull a renewal date and contract identifier from one document, retain the supporting text, and verify the date's format. Keep the model away from deciding which commercial terms override others until your evaluation covers those exceptions.

Operators: Shorten the Work Before Shortening the Budget

For a mid-market CTO, retrieval can matter more than the headline context limit. Give a summarizer the relevant meeting or policy section, then evaluate whether it preserves decisions, owners and exceptions. A summary that quietly drops an obligation creates work elsewhere in the organization.

A senior operator choosing a routing system should measure misrouted tickets, unresolved cases and time to an accepted decision. Define escalation from observable evidence: missing required fields, conflicting source passages, an unsupported category or a failed check. A model's own confidence statement is an insufficient routing policy on its own.

Buyers: Specify the Unit You Want to Pay For

Ask a supplier to quote cost per accepted extraction or completed case, with token counts, effort, retries and review included. The $60 classifier budget says little about a service that repeatedly retries failed requests or sends a large policy book with every ticket.

For subagents, meaning models assigned smaller tasks inside a larger workflow, give Haiku a bounded evidence-gathering job. It can be the candidate that finds a relevant passage or returns a compact summary while Sonnet handles the broader decision. Evaluate the handoff: the larger model needs enough source material to check the subagent's work.

Claude Haiku vs Sonnet: The Decision Rule

Choose Haiku when a narrow answer is easy to verify; retain Sonnet when judgment, coordination or recovery determines whether the job succeeds. Sonnet 5.5's fresh-token rates are 20 times Haiku's short-prompt rates and four times its long-prompt rates. The premium needs to buy better accepted outcomes.

Anthropic's release table supports evaluating the new Haiku, but gives a reason to keep Sonnet on harder coding. The following are Anthropic's published benchmark results, with Haiku 5.5, Haiku 4.5 and OpenAI GPT-6 Luna in that order:

  • GDPval-AA v2.1, knowledge work: 1620, 735 and 1437.
  • OSWorld 2.1, offline subset, computer use: 72.4%, 15.7% and 48.9%.
  • Terminal-Bench 4.0, agentic coding: 39.2%, 0.0% and 16.4%. Anthropic reports 70.6% for Sonnet 5.5 on this benchmark.

Source: Anthropic's performance table. Benchmark scores answer questions about those evaluation tasks. They do not establish an acceptance rate for your contracts, tickets or codebase.

Keep Sonnet as the lead for a repository-wide change that must discover dependencies, coordinate edits and recover from failed steps. Consider Haiku for a subtask such as summarizing one file or extracting a specific test failure. Anthropic itself recommends the larger models for complex agentic coding in the announcement.

Validate First, Then Pay for Escalation

An illustrative mixed workflow runs every classification on Haiku and sends failed validations or unresolved cases to Sonnet. If 10% need one additional Sonnet call, with the same token counts as the baseline, the monthly bill is $60 + $120 = $180. That is $1,020 less than running all calls on Sonnet. The example excludes validation infrastructure and assumes the second call needs no extra context.

The first Haiku call is still billed on every escalated case. A forecast that prices only the final successful model misses that cost. Measure retries as attempts per completed job, and compare end-to-end time as well as tokens.

A clay library workflow sends a ticket from Haiku to a validation desk, then passing work to Done and review cases upstairs to Sonnet before rejoining Done
Your application validates and escalates. Both model calls count when a case needs review.

Use the Sonnet 5.5 setup guide for the larger-model path. For budgeting, apply the current $0.10 per MTok cache-read rate verified here; Anthropic cut it from $0.20 on October 7. Price change.

Claude Haiku 4.5 API: Moving to 5.5

Treat the switch as an integration change as well as a model-name change. A service that worked with Haiku 4.5 can fail because of request parameters, response parsing or output limits even when the new model answers the task well.

Work through Anthropic's Haiku migration guide for your platform. The consequential changes are:

  1. Update the ID and recount. Use claude-haiku-5-5 on the Claude API. Recount prompts with that model and revisit cost estimates and output caps.
  2. Update thinking and sampling settings. Replace manual thinking.type: enabled and budget_tokens with adaptive thinking. Set output_config.effort for the workload. Remove temperature, top_p and top_k as instructed by the guide.
  3. Remove assistant prefill. A final assistant message supplied for the model to continue is rejected. End the request with a user turn and choose the supported output-format mechanism for your platform.
  4. Read response blocks by type. A response can begin with a thinking block. Select the text or tool-use blocks you need, and handle output exhaustion before a visible answer.
  5. Handle refusals and conversation replay. Route a stop_reason of refusal to an application path. If replaying thinking, preserve the blocks unchanged, use the producing or linked account, and follow the guide's unchanged-prefix rules.

For a CTO with a Priority Tier capacity commitment, there is a separate reason to wait: Anthropic's migration guide says Haiku 5.5 does not support Priority Tier. Establish an acceptable capacity path before moving the queue. Cheap tokens do not replace a service requirement. Capacity limitation.

Haiku Retirement Dates: Read the Status, Not Just the Date

On 8 October 2026, Anthropic's deprecations page lists Haiku 4.5 as Active, with retirement not sooner than October 15, 2026. That is a lower bound, not an announced shutdown on that date. It lists Haiku 5.5 as Active, with retirement not sooner than October 7, 2027. Claude API model status.

For older integrations, Haiku 3.5 retired on the Claude API on February 19, 2026, and Haiku 3 on April 20, 2026. Those are retirement dates on the API lifecycle page, rather than a schedule inferred from the 5.5 release. Deprecation history.

Plan a Haiku 4.5 migration around evaluation and integration readiness. The current status gives you room to make that decision deliberately; continue checking the official lifecycle page for an announced retirement.

What Is Overhyped

Three attractive claims need qualification before they become a budget or architecture decision.

“90% cheaper” describes the short-prompt rate difference. It does not guarantee a 90% reduction for the same completed job. Tokenization, thinking, request length, repeated attempts and escalations affect the bill. The equal-count example isolates pricing; your observed counts determine the forecast.

A million-token window does not make long prompts inexpensive. Haiku's upper tier begins well before that capacity limit. A summarizer that carries irrelevant history into every request can spend more without returning a better summary. Count the prompt after assembling it, and retain enough room for the output you need.

An effort dial does not turn every Haiku task into a Sonnet task. Raising effort is worth evaluating when it fixes an observed failure. Keep Sonnet in that evaluation when the problem is sustained coding or difficult judgment. Compare completed outcomes and elapsed time, including retries.

Sonnet's cache-read cut also changes the comparison for an existing agent. Recompute its current bill before deciding to migrate: repeatedly reused input is now cheaper on Sonnet, while fresh input and output keep their list rates. The Claude API pricing guide puts this release inside the wider model and feature budget.

What to Do This Week

Act now on a narrow workload with a clear acceptance check. Wait on a queue whose quality or capacity requirements you have not resolved. A founder paying for repeated classifications has an immediate evaluation candidate; a CTO depending on Priority Tier needs a capacity decision first. A service already meeting its budget and quality target can keep its current model while the comparison runs.

  1. Pick one queue and define a passing result

    Use representative inputs with known correct answers, including ambiguous and difficult cases. Specify which errors require review and which failures can safely be retried. Start with classification, extraction or a constrained summary.

  2. Measure the same work on each candidate

    Run Haiku 4.5, Haiku 5.5 and Sonnet 5.5 on those inputs with suitable effort settings. Record model-specific input counts, billed output, completed-job time and accepted results. Include every attempt in the cost calculation. Use the new tokenizer's counts when checking the prompt-length tier.

  3. Move the passing slice and keep an escalation path

    Complete the migration checks, then move the subset whose observed quality and cost satisfy your requirements. Validate outputs, inspect failures and retain Sonnet for the harder work. A successful result is a cheaper accepted decision at your required speed.

Get future model and API budget updates in the newsletter.

Published
Category
AI
Related Articles
Otter AI Alternatives: Granola, Krisp, MeetGeek, Circleback and Claap

Otter AI Alternatives: Granola, Krisp, MeetGeek, Circleback and Claap

Choose Granola, Krisp, MeetGeek, Circleback or Claap by the job. Compare live prices, free limits, bot-free capture and CRM handoffs.Oct 8, 2026AI
Mistral Large 4: Pricing, Open Weights, and When to Use It

Mistral Large 4: Pricing, Open Weights, and When to Use It

Mistral Large 4: verified API prices, the launch offer, claimed benchmarks, and when to test the preview or wait for self-hosting.Oct 7, 2026AI
Lindy AI Review (Verified October 2026)

Lindy AI Review (Verified October 2026)

Lindy AI review: current plans, credit costs, email and meeting setup, integration limits, and when Gumloop, n8n or Marblism fits better.Oct 7, 2026AI
Mistral Pricing (2026): API Costs and Vibe Seats

Mistral Pricing (2026): API Costs and Vibe Seats

Mistral API rates, Le Chat and Vibe plans, free limits, EU hosting costs and a worked monthly bill. Verified October 2026.Oct 6, 2026AI
Claude Skills

Claude Skills

Use Claude Skills across the apps, Claude Code and API. Build a weekly client-report skill file by file, then install and share it.Oct 6, 2026AI
AI Employee: What It Does, Costs, and How to Start

AI Employee: What It Does, Costs, and How to Start

What AI employees can do, current prices for six tools, and how to pick one role, keep human review and measure a 30-day pilot.Oct 6, 2026AI
Wispr Flow Alternatives (2026): VoiceOS, Superwhisper, Aqua Voice, Willow and Monologue

Wispr Flow Alternatives (2026): VoiceOS, Superwhisper, Aqua Voice, Willow and Monologue

Compare VoiceOS, Superwhisper, Aqua Voice, Willow and Monologue by work task, current price, offline privacy, team controls and coding-agent input.Oct 5, 2026AI
AI Agent Examples: 13 Small Business Jobs and What They Cost

AI Agent Examples: 13 Small Business Jobs and What They Cost

13 AI agent examples for small businesses: what runs the job, what a person checks, and current prices from Tidio, Vida, MindStudio and more.Oct 5, 2026AI
Newsletter

One letter, every Sunday.Working systems, not hot takes.

Weekly. No spam. Unsubscribe anytime.