The biggest change is price. Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens for requests under 100,000 tokens.
Haiku 4.5 was priced at $1 per million input tokens and $5 per million output tokens, regardless of request size. For larger Haiku 5.5 requests, Anthropic charges $0.50 per million input tokens and $2.50 per million output tokens.
Anthropic says roughly 90% of Haiku 4.5 requests fell into the lower-length category now covered by the cheapest Haiku 5.5 pricing. The company estimates that customers will save about 75% on average, taking into account request mix and changes to its tokenizer.
The new model also introduces effort controls to the Haiku line for the first time. Developers can decide how many tokens the model should spend on a task, with medium set as the default.
Anthropic’s own evaluations show large gains over Haiku 4.5. On the OSWorld 2.1 offline subset for computer use, Haiku 5.5 scored 72.4%, compared with 15.7% for Haiku 4.5 and 48.9% for OpenAI’s GPT-6 Luna.
On the GDPval-AA v2.1 knowledge-work benchmark, Haiku 5.5 scored 1,620, compared with 735 for Haiku 4.5 and 1,437 for GPT-6 Luna. Sonnet 5.5 remained ahead with a score of 1,840.
The gap was also visible in agentic coding. On Terminal-Bench 4.0, Haiku 5.5 scored 39.2%, while Haiku 4.5 scored 0% and GPT-6 Luna scored 16.4%. Sonnet 5.5 led that test at 70.6%.
Anthropic is extending Haiku beyond the summarization, classification and routing workloads typically associated with smaller models. The company says Haiku 5.5 can also handle database queries, compaction and agentic tasks where latency matters, including browser use and live customer support.
The model also arrives with beta support for computer and browser use in Anthropic’s Python and TypeScript SDKs.
Haiku 5.5 faces competition from other small models that are also targeting low-cost inference. Anthropic is directly comparing the model with GPT-6 Luna, while developers are also evaluating alternatives from companies including Z.ai and Alibaba.
Artificial Analysis reports that Z.ai’s GLM-5.3-Flash scores 1,647 on GDPval-AA v2.1 and 1,454 on AA-Briefcase v1.1. Anthropic reports scores of 1,620 and 1,578, respectively, for Haiku 5.5.
Alibaba’s Qwen3.7 Flash is cheaper on some workloads. Its international pricing starts at $0.03 per million input tokens and $0.13 per million output tokens for inputs up to 32,000 tokens.
Anthropic is making additional pricing changes alongside the Haiku launch. The company is cutting Sonnet 5.5’s cache-read price from $0.20 to $0.10 per million tokens and says the reduction should lower the cost of most agentic tasks by about 20%.
The company is also adding monthly Claude Platform API credits to Max and Team subscriptions. Max 5x subscribers will receive $100 per month, Max 20x subscribers will get $200, and Team plans will receive as much as $500 pooled across users.
Haiku 5.5 also includes stricter cybersecurity protections than Haiku 4.5, although Anthropic says the model still permits a wider range of defensive work than Sonnet 5.5. Penetration testing remains blocked, while organizations that need broader access for cybersecurity or biology workloads can apply through Anthropic’s verification programs.
This analysis is based on reporting from Yahoo Finance.
Image courtesy of Anthropic.
This article was generated with AI assistance and reviewed for accuracy and quality.