Anthropic Unveils Claude Haiku 5.5: 75% Cheaper, Faster, and Smarter Small Model
Anthropic's Small Model Gets a Major Upgrade
Anthropic has officially released Claude Haiku 5.5, the newest addition to its model family, positioning it as the cheapest, fastest, and most capable small model the company has ever shipped. The launch is a significant move for developers and enterprises looking to scale AI workloads without breaking the bank, especially for high-volume, repetitive tasks that were previously cost-prohibitive.
The new model is available immediately across all major platforms, including Amazon Bedrock, Google Cloud, Microsoft Azure, and directly through the Claude API. Developers can start using it today with the model ID claude-haiku-5-5.
Pricing: A 75% Cost Reduction
The headline feature of Haiku 5.5 is its aggressive pricing. According to Anthropic, the model costs around 75% less to run than its predecessor, Haiku 4.5, on average. For requests up to 100,000 tokens—which constitute roughly 90% of all usage on the previous Haiku model—the price drop is even steeper at 90%.
- Input tokens: $0.10 per million (vs. $1.00 for Haiku 4.5)
- Output tokens: $0.50 per million (vs. $5.00)
- Cache reads: $0.01 per million (vs. $0.10)
- Cache writes: $0.125 per million (vs. $1.25)
This pricing puts Haiku 5.5 in direct competition with other compact models in the market, making it an attractive option for startups and large enterprises alike that need to process millions of requests daily.
Performance: Benchmark Gains Across the Board
Claude Haiku 5.5 isn't just cheaper—it's dramatically more capable. Anthropic's published benchmarks show massive improvements over Haiku 4.5, and in some cases, it even rivals or surpasses much larger models.
On the GDPval-AA v2.1 knowledge work benchmark, Haiku 5.5 scored 1620, more than double Haiku 4.5's 735, and even beating OpenAI's GPT-6 Luna (1437). Similarly, on the AA-Briefcase v1.1 test, it scored 1578 versus 614 for the previous generation.
In computer use (OSWorld 2.1), Haiku 5.5 achieved 72.4% accuracy on the offline subset, a massive leap from Haiku 4.5's 15.7%. This makes it particularly well-suited for browser automation and agentic workflows.
For coding, the model shows strong results on Terminal-Bench 4.0 (39.2% vs. 0.0% for Haiku 4.5) and FrontierCode 1.1 (46.4%), though it still trails the larger Sonnet 5.5 (70.6% and 52.1%, respectively). Anthropic recommends using Haiku 5.5 as a subagent for coding tasks, paired with Opus or Sonnet as the lead model.
First Haiku With Adjustable Effort
One of the most notable technical additions is the adjustable effort setting, a feature previously reserved for Anthropic's larger models. This allows developers to fine-tune the model's reasoning depth based on their cost and latency requirements.
The effort setting directly impacts performance. At low effort, Haiku 5.5 is extremely fast and cheap but less accurate; at high effort, it approaches the reasoning capabilities of much larger models. This flexibility is a game-changer for production environments where different tasks require different levels of intelligence.
Real-World Customer Results
Several enterprise customers have already tested Haiku 5.5 in production environments, with impressive results:
- Asana reported a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn in their AI Teammates product.
- HubSpot saw the best CRM evaluation score they've ever recorded (92.8%) on simulated portal tests, with the fastest completion times and lowest false-positive rates.
- AlphaSense noted a statistically significant improvement over Haiku 4.5 (0.84 vs. 0.76) on their document Q&A workload, which handles about 8 million calls per week.
- Box measured an 11-point improvement over Haiku 4.5 at roughly half the latency for analytical tasks like cost reports and financial summaries.
- Cognition integrated Haiku 5.5 into Devin Fusion as a "sidekick," achieving a top-tier FrontierCode score of 66.2 while cutting cost and latency.
Sonnet 5.5 Price Cut and New API Credits
In addition to the Haiku launch, Anthropic is making its broader ecosystem more affordable. The company has halved the price of cache reads on Claude Sonnet 5.5, dropping from $0.20 to $0.10 per million tokens. Since cache reads make up a large share of token consumption in agentic workflows, this reduces the overall cost of Sonnet 5.5 by approximately 20% on most tasks.
Anthropic is also introducing a monthly API credit for Max and Team subscribers. Max 5x users receive $100 per month, Max 20x users get $200, and Team subscribers get up to $500 pooled across users. These credits are designed to encourage experimentation with building agents and applications on the Claude Platform.
Safety and Alignment Improvements
Haiku 5.5 comes with significant safety upgrades. Anthropic reports major improvements across almost all alignment evaluations relative to Haiku 4.5, including fewer instances of misaligned behavior and lower willingness to cooperate with misuse.
Cybersecurity safeguards are more restrictive than Haiku 4.5's but somewhat less restrictive than those applied to Sonnet 5.5, permitting a wider range of defensive tasks while still blocking penetration testing and other attack techniques. Biology safeguards match the standards set for Sonnet 5, Sonnet 5.5, and Opus 5.
Technical Specifications and Availability
Claude Haiku 5.5 supports a 1 million-token context window and up to 128K output tokens, making it suitable for long-document processing and complex multi-turn interactions. It is multimodal, accepting both text and image inputs, and supports native PDF input, tool calling, and structured output.
The model is available on Amazon Bedrock, Google Cloud, Microsoft Azure, and the Claude Platform on AWS, with regional data residency options on AWS. Developers can also access it through Anthropic's updated Python and TypeScript SDKs, which now include beta support for computer use and browser use.
Why It Matters
The launch of Claude Haiku 5.5 represents a significant shift in the economics of AI deployment. By offering near-flagship performance at a fraction of the cost, Anthropic is enabling a new class of applications that were previously impractical—such as real-time customer support, high-frequency classification, and large-scale document processing.
For enterprises, the combination of lower pricing, adjustable effort, and improved safety makes Haiku 5.5 a compelling option for production workloads. As the AI model race intensifies, the focus is increasingly shifting from raw capability to cost-efficiency and practical deployment—and Haiku 5.5 appears to be leading that charge.
Related News

Talorys: Self-Hosted AI Agent on Cloudflare's Free Tier

Bitwarden's Dual License Shift Stirs Open Source Community Concerns

Lean Theorem Prover Faces Reliability Scrutiny Amid AI Math Surge

TypeSafe AI Raises $870M to Build Machine-Native Models

OpenAI Withdraws Three Math Preprints After Sign Error in AI Proofs

