On October 7, 2026, Anthropic launched Claude Haiku 5.5, the fastest and lowest-cost small language model within the Claude family. The model brings a 90% reduction to both API input and output prices compared to its predecessor Haiku 4.5. Input pricing dropped from $1.00 to $0.10 per million tokens, while output pricing fell from $5.00 to $0.50 per million tokens. For the first time, Anthropic introduced adjustable effort parameters to control reasoning intensity. Its score on the OSWorld 2.1 computer-use benchmark jumped sharply from 15.7% to 72.4%, and the context window has been expanded to 1 million tokens. This article is built upon Anthropic’s official release and pricing documentation. It covers tiered pricing tables, a model selection matrix against Sonnet 5.5, three distinct cost-saving strategies, and ready-to-use API code examples to help developers finalize model selection and cost planning.
Benchmark Performance: Comparative Scores Against Haiku 4.5, GPT-6 Luna and Sonnet 5.5
The table below presents the benchmark results across core evaluation tracks, comparing Haiku 5.5 with Haiku 4.5, GPT-6 Luna and Sonnet 5.5 for reference. All data comes from the official evaluation suite published by Anthropic.
| Benchmark Category | Benchmark Name | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|---|
| Knowledge work | GDPval-AA v2.1 | 1620 | 735 | 1437 | 1840 |
| Knowledge work | AA-Briefcase v1.1 | 1578 | 614 | 1336 | 1824 |
| Computer use | OSWorld 2.1 (Offline subset) | 72.4% | 15.7% | 48.9% | 83.9% |
| Multidisciplinary reasoning | Humanity’s Last Exam (no tools) | 45.9% | 10.2% | — | 56.9% |
| Multidisciplinary reasoning | Humanity’s Last Exam (with tools) | 57.4% | 18.7% | — | 64.5% |
| Agentic coding | Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% |
| Agentic coding | FrontierCode 1.1 (Main) | 46.4% | — | 42.4% | 52.1% (xhigh) |
| Visual reasoning | Chartography (no tools) | 46.4% | 6.4% | 29.1% | 61.6% |
The most striking improvements are summarized separately in the table below, showing the magnitude of upgrades from Haiku 4.5:
| Benchmark | Haiku 5.5 | Haiku 4.5 | Improvement |
|---|---|---|---|
| GDPval-AA v2.1 (Elo score) | 1620 | 735 | +120% |
| OSWorld 2.1 Computer Use Accuracy | 72.4% | 15.7% | +56.7 percentage points |
| Humanity’s Last Exam (No tools) | 45.9% | 10.2% | +35.7 percentage points |
| Terminal-Bench 4.0 | 39.2% | 0.0% | New capability |
Anthropic also released real-world production metrics shared by enterprise beta customers, demonstrating practical performance gains in business workflows:
- Asana: Task completion latency reduced by 30%, single-turn reasoning speed increased by 2.5 times.
- HubSpot: 92.8% accuracy on CRM task suites, marking the highest historical accuracy for the product.
- AlphaSense: Document QA accuracy reached 0.84, an improvement over Haiku 4.5’s 0.76.
- Box: Composite evaluation score 11 points higher than Haiku 4.5, with latency cut in half.
These results show that Haiku 5.5 is not merely a price reduction update. It closes major capability gaps in computer automation, agentic coding and visual reasoning. The previous Haiku 4.5 was limited to simple classification and extraction jobs. Haiku 5.5 can now reliably handle browser and desktop automation tasks and act as a sub-agent inside multi-agent systems.
Pricing Breakdown: Tiered Billing and Three Cost Reduction Pathways
Claude Haiku 5.5 uses a two-tier pricing model. Requests under and over the 100,000-token threshold are billed separately, effective October 2026, based on Anthropic official pricing pages.
| Billing Dimension | ≤100K Tokens per Request | >100K Tokens per Request |
|---|---|---|
| Input | $0.10 / MToken | $0.50 / MToken |
| Output | $0.50 / MToken | $2.50 / MToken |
| Cache write | $0.125 / MToken | $0.625 / MToken |
| Cache read | $0.01 / MToken | $0.05 / MToken |
The prior generation Haiku 4.5 applied flat pricing: $1.00 input and $5.00 output per million tokens, with no tier split. For requests within the 100K token cap, the new pricing delivers exactly a 90% price cut. Even for prompts exceeding 100K tokens, users still see approximately 50% savings versus Haiku 4.5. The Batch API for bulk processing adds an extra 50% discount that stacks with other eligible savings.
Anthropic has introduced three independent optimization pathways for teams to further lower inference expenses.
1. Prompt Caching
Developers can cache repeated system prompts or long static document prefixes. Reading from cache reduces the cost of cached input down to 10% of standard input pricing, at $0.01 per million tokens. This method can deliver up to 90% savings for workloads with high prompt reuse, such as document parsing pipelines, standardized classification jobs, and repeated agent system instructions.
2. Batch API
Asynchronous tasks that do not require real-time responses can be submitted via the Batch API endpoint. All jobs submitted to this interface automatically receive a 50% discount. Batch processing is ideal for offline bulk tasks: mass document extraction, dataset labeling, backlog content summarization and bulk data enrichment.
3. Effort Parameter Tuning
Haiku 5.5 is the first model in the Haiku family with configurable effort controls. For low-complexity tasks, developers can set effort: "low". This setting shortens internal reasoning chains and cuts down token consumption. The default value for this parameter is medium. When higher accuracy is mandatory, users can switch it to high to enable deeper reasoning at the cost of slightly higher token usage.
Model Selection: Haiku 5.5 vs Sonnet 5.5
Both models support a 1 million-token context window, a maximum output limit of 128K tokens, text and image multimodal input. Their knowledge cutoff date is June 2026. Sonnet 5.5 is positioned as the balance point between speed and intelligence. Its base input price is $2.00 per million tokens, 20 times the base input price of Haiku 5.5. Haiku 5.5 targets high-volume, low-latency workloads including classification, information extraction and routing.
Suitable use cases for Claude Haiku 5.5
- High-throughput text classification, summary generation and entity extraction pipelines
- Real-time customer service dialogue and low-latency voice assistant turn responses
- Sub-agent execution inside multi-model agent architectures, for subtasks delegated by larger foundation models
- Form filling, data entry and browser or desktop automation workflows
Not recommended scenarios for Haiku 5.5
- Complex analytical reports requiring multi-step deep reasoning
- Large-scale code refactoring and system architecture design
- Legal, medical compliance content where near-perfect output accuracy is required
Haiku 5.5 defaults to effort: "medium", while Sonnet 5.5 defaults to high. If Haiku 5.5 running under effort: "high" still fails to meet required accuracy thresholds, developers should migrate the workload to Sonnet 5.5 instead of continuing to tune Haiku.
Quick Integration and Python SDK Code Example
The model ID for Haiku 5.5 is claude-haiku-5-5. Anthropic states the retirement date will be no earlier than October 7, 2027. This identical model ID works across AWS Bedrock (anthropic.claude-haiku-5-5), Google Cloud Vertex AI and Microsoft Foundry with no extra configuration changes.
The code sample below uses the official Anthropic Python SDK:
Developers who want to run cross-model benchmarking and A/B testing across multiple LLM providers can use a unified platform for standardized API access. 4sapi, functioning as an API gateway, allows developers to switch between many mainstream large language models including the full Claude lineup using a single API key, with interfaces following standard OpenAI SDK conventions. This removes repetitive integration work when comparing Haiku against other competing models in production test environments.
FAQ
Q: Should I choose Claude API pay-as-you-go billing or a Claude.ai Pro subscription?
These two products target different user groups. The Claude.ai Pro subscription is built for individual daily chat use, and it does not grant API access. The pay-as-you-go API plan serves developers and enterprises building AI features into products and services. It has no seat limits, charges purely based on consumption, and supports mixing multiple model families within one application. The API route is the only viable path for embedding Claude capabilities into external software. A subscription is simpler for personal casual chat use. Full pricing details can be checked on Anthropic’s official pricing page.
Closing Remarks
Claude Haiku 5.5 marks a major upgrade for Anthropic’s lightweight model tier. The combination of drastically reduced pricing, 1M-token context window, adjustable reasoning effort and vastly improved computer-use performance redefines what developers can build with low-cost high-throughput models.
The tiered pricing structure rewards workloads that fit within the 100K token per-request limit, but teams must carefully evaluate prompt lengths to avoid unexpected cost jumps once crossing that threshold. The three cost-reduction mechanisms — prompt caching, batch jobs and effort tuning — can be stacked to further cut inference spend for suitable workloads.
For production planning, Haiku 5.5 works best as a routing, extraction and automation workhorse. When tasks demand deep multi-step reasoning or strict compliance accuracy, Sonnet 5.5 remains the better fit. All data in this article is sourced from Anthropic’s October 7 official release announcement and public pricing documentation.
International access: https://4sapi.com
Domestic access: https://4sapi.cn




