Key takeaways
- Anthropic launched Claude Haiku 5.5 on October 7, 2026. The Claude API ID is claude-haiku-5-5; Anthropic documents a 1-million-token context window, up to 128K output tokens, adaptive thinking and adjustable effort.
- Direct Claude Platform pricing is tiered by prompt length: up to 100K prompt tokens costs $0.10 input and $0.50 output per million tokens; over 100K costs $0.50 input and $2.50 output. Cache reads are $0.01 or $0.05 and cache writes are $0.125 or $0.625 across those tiers.
- The model is available on the Claude Platform, Claude.ai, Claude Code, AWS, Google Cloud and Microsoft Foundry. AWS examples use global.anthropic.claude-haiku-5-5, with additional US, EU, AU and JP Geo CRIS profiles documented.
- Anthropic reports major gains over Haiku 4.5 and lower average task cost, but those results are vendor-run. The company also says its newer tokenizer makes the same text count as roughly 30% more tokens than on Haiku 4.5.
- Test on complete production-shaped tasks before replacing a larger model or Haiku 4.5. Record prompt tier, token counts, effort, quality, latency, retries, tool behavior and cost per accepted result.
Claude Haiku 5.5 launched on October 7
Anthropic introduced Claude Haiku 5.5 on October 7, 2026 as the cheapest and fastest member of the Claude 5.5 family. The company targets high-volume, latency-sensitive work: classification, extraction, routing, summarization, compaction, live support, narrowly scoped coding and subagent tasks. The Claude Platform model ID is claude-haiku-5-5.
The live documentation lists a 1-million-token context window, up to 128K output tokens, text and image input, text output, tool use and adaptive thinking with an effort control. Those are interface capabilities, not evidence that every task benefits from million-token prompts or that a small model can safely inherit broad agent authority.
The $0.10 rate ends when a prompt exceeds 100K tokens
Haiku 5.5 has two direct Claude Platform price tiers. For prompts up to 100,000 tokens, Anthropic lists $0.10 per million input tokens, $0.50 per million output tokens, $0.01 for cache reads and $0.125 for cache writes. For prompts over 100,000 tokens, the rates rise to $0.50 input, $2.50 output, $0.05 cache reads and $0.625 cache writes per million tokens.
That threshold matters for routing and budgeting: one long prompt moves the request onto the higher tariff. Anthropic says prompts within 100K represented about 90% of requests to the previous Haiku model, and says Haiku 5.5 costs around 75% less to run on average. The average incorporates its traffic mix and token-use assumptions; it is not a guaranteed reduction for a specific application.
A new tokenizer changes migration economics
Anthropic’s model documentation says Haiku 5.5 uses the newer tokenizer used by Claude 4.7 and later models, so the same text counts as approximately 30% more tokens than on Haiku 4.5. Anthropic’s pricing footnote says this tokenization change is included in its average-cost comparison. Teams should therefore avoid estimating savings by multiplying old Haiku 4.5 token logs by the new list rate alone.
Replay actual prompts and responses through the candidate route, including retrieval context, system instructions, tool results, thinking, retries and cache behavior. Split results by prompts below and above 100K. The useful metric is cost per accepted task at the required quality and latency—not the lowest advertised input-token price.
Anthropic reports broad gains, but larger models still lead hard coding tasks
Anthropic reports Haiku 5.5 scores of 72.4% on the OSWorld 2.1 offline subset, 39.2% on Terminal-Bench 4.0 and 46.4% on FrontierCode 1.1 Main. It also reports higher scores than Haiku 4.5 across the launch table. These are vendor-selected, vendor-run results under the linked system card and benchmark footnotes, not independent AccessAllGPT measurements.
The same launch shows Sonnet 5.5 at 70.6% on Terminal-Bench 4.0 and explicitly says Sonnet and Opus remain better choices for complex agentic coding. Haiku’s case is narrower: cheap, frequent, well-bounded work where speed and volume matter. Do not turn a benchmark improvement into permission for unattended browser, desktop or coding actions.
Availability spans Anthropic and three cloud platforms
Anthropic says Haiku 5.5 is available now on the Claude Platform, Claude.ai, Claude Code, Amazon Web Services, Google Cloud and Microsoft Foundry. Its consumer model page says Free, Pro, Max, Team and Enterprise users can select it on web, iOS and Android. Availability still depends on the account, product, region and administrator controls.
AWS documents global.anthropic.claude-haiku-5-5 for Global CRIS and says US, EU, AU and JP Geo CRIS profiles are also available on bedrock-runtime. In AWS GovCloud (US), it lists both bedrock-runtime and bedrock-mantle. The AWS code warns that a response may place a thinking block before the text block, so integrations should select content by block type rather than assuming text is at a fixed array index.
Safety claims remain provider evidence
Anthropic says Haiku 5.5 improves on almost all of its alignment evaluations versus Haiku 4.5, with fewer instances of misaligned behavior and less willingness to cooperate with misuse. It says cyber safeguards are stricter than Haiku 4.5 but less restrictive than those on Sonnet 5.5, while biology safeguards match Sonnet 5, Sonnet 5.5 and Opus 5. The linked system card documents the provider’s evaluation work.
Those findings do not certify a particular application. A cheaper model may be invoked more often and across more workflows, increasing aggregate exposure. Keep deterministic authorization, least-privilege tools, budget limits, logging and review around agent actions, and test both unsafe compliance and harmful over-refusal on the actual provider route.
What builders should do next
Use a shadow evaluation before changing traffic. Freeze representative short and long prompts, expected outputs, tool schemas and acceptance criteria. Compare the current model with claude-haiku-5-5 at each effort level you might deploy. Record prompt length, tier, input, output and cache tokens, first-token and completion latency, retries, parsing errors, tool-call validity, refusals and cost per accepted result.
Adopt Haiku 5.5 for bounded high-volume work when quality and cost gates pass. Constrain it behind review where browser, desktop or code tools can write. Wait when account access, cloud routing, data location or the over-100K tariff is unclear. Reject a migration that relies on headline pricing while ignoring tokenizer growth, tier crossings, retries or weaker outcomes on the tasks that matter.
Copy-ready Haiku 5.5 rollout record
Complete one record per provider route, workload and prompt-length band before shifting production traffic.
Entries stay in this browser tab and are not submitted to AccessAllGPT. Blank responses are copied as [Unresolved].
Task, owner, users, volume, consequence, trial or deploy decision, review date and expiry.
Claude Platform, AWS, Google Cloud or Microsoft Foundry; exact model ID, account, region/profile, SDK and API version.
Observed prompt tokens, up-to-100K or over-100K rate, expected tier crossings and maximum context policy.
Effort, thinking, max output, temperature, system prompt, tools, cache policy, timeout and retry ceiling.
Tokenizer comparison, thinking-block parsing, tool calls, long prompts, truncation, refusals and malformed responses.
Frozen tasks, scorer, baseline, acceptance rate, first-token and completion latency, variance and failure rate.
Input, output, cache write/read and thinking tokens; attempts, retries, provider charges and cost per accepted task.
Readable data, callable tools, write permissions, approvals, monitoring, spend stop, rollback model and re-evaluation trigger.
Primary sources
Browse the publication-wide evidence index →
- Introducing Claude Haiku 5.5Anthropic · Reviewed: October 7 announcement date; positioning and use cases; benchmark table; effort control; customer reports; tiered token pricing; safety; platform availability; Sonnet 5.5 cache-price change; API credits; footnotes · Retrieved · Supports: Anthropic launched Claude Haiku 5.5 on October 7, 2026, identified claude-haiku-5-5 as the Claude Platform model ID, published two prompt-length price tiers, and said the model is available through its platform and major cloud providers.
- Claude Haiku 5.5 model overviewClaude Platform Docs · Reviewed: Model positioning; Claude API ID; 1M-token context; 128K-token maximum output; adaptive thinking and effort; tiered pricing; newer tokenizer; account-bound thinking blocks; platform IDs; lifecycle information · Retrieved · Supports: Anthropic documents Haiku 5.5 with a 1-million-token context window, up to 128K output tokens, adaptive thinking, configurable effort and a newer tokenizer that counts the same text as approximately 30% more tokens than Haiku 4.5.
- Claude Haiku 5.5 System CardAnthropic · Reviewed: Release system-card document linked from the benchmark and safety sections; evaluation scope; alignment, misuse, cyber and biology safety coverage; limitations of model evaluation · Retrieved · Supports: Anthropic provides the release system card as its detailed first-party record for Haiku 5.5 evaluation methods and safety findings; the launch page summarizes those findings as vendor evidence rather than independent certification.
- Introducing Claude Haiku 5.5 on AWSAmazon Web Services · Reviewed: October 7 availability announcement; model positioning; prerequisites; Boto3 InvokeModel and Converse examples; Anthropic SDK route; thinking-block response handling; CRIS profiles; GovCloud endpoints; regional and pricing links · Retrieved · Supports: AWS says Haiku 5.5 is available on Amazon Bedrock through US, EU, AU, JP and Global CRIS profiles, documents global.anthropic.claude-haiku-5-5 in examples, and lists Bedrock Runtime and GovCloud access routes.
Limitations
AccessAllGPT reviewed four first-party Anthropic and AWS surfaces but did not receive Claude, AWS, Google Cloud or Microsoft credentials for this work. We did not call Claude Haiku 5.5; inspect account-specific availability or prices; reproduce OSWorld, Terminal-Bench, FrontierCode, GDPval-AA, Humanity’s Last Exam, Chartography or customer evaluations; compare Haiku 4.5, GPT-6 Luna or Sonnet 5.5; measure tokenizer changes, prompt-tier crossings, context reliability, latency, token use, cache behavior or cost per task; test adaptive effort, thinking blocks, browser or computer use; audit alignment, cyber or biology safeguards; inspect training data; or perform a security, privacy, residency or accessibility assessment. Provider IDs, routes, limits, prices, credits, policies and availability can change.
Disclosures
AccessAllGPT received no Anthropic or AWS credentials, credits, early access, briefing, benchmark data, review, payment or compensation for this article. Anthropic and Amazon Web Services did not sponsor, review or endorse it. AccessAllGPT did not use, test, score or rank Claude Haiku 5.5. AccessAllGPT Research is operated by NeuralArc, is independent, and is not affiliated with Anthropic, AWS or organizations cited. Publication-wide relationships are listed on the disclosures page.
Further AccessAllGPT guidance
- Anthropic Launches Claude Sonnet 5.5
- OpenAI Rolls Out GPT-6 and Intelligent UI
- OpenAI Publishes a GPT-6 Family Guide
- Move an LLM API Without Breaking Production
- Choose a Model Without Chasing the Leaderboard
- AI API Data Retention and Residency
- AccessAllGPT Research methodology
- Evidence standards
- Publication disclosures
Continue the research
Get evidence-led updates for teams making production AI decisions.