Anthropic released Claude Haiku 5.5 on Wednesday, positioning it as the company’s cheapest, fastest, and most capable small model. The new pricing structure sets costs at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, a 90% reduction from the previous Haiku 4.5 rates of $1 and $5 respectively. This pricing matches that of OpenAI’s GPT-6 Luna, which launched in late September. Anthropic estimates the average saving for users is near 75%, given that approximately 90% of requests to the prior model fell under the 100,000-token threshold.
In benchmark tests, Haiku 5.5 scored 72.4% on the OSWorld 2.1 offline subset and 39.2% on Terminal-Bench 4.0, surpassing GPT-6 Luna’s scores of 48.9% and 16.4%. However, it trails Anthropic’s larger Sonnet 5.5 model, which scored 70.6% on coding tasks. On GDPval-AA v2.1, Haiku 5.5 achieved an Elo rating of 1620, compared to Luna’s 1437 and Haiku 4.5’s 735. The release follows Opus 5.5 and Sonnet 5.5, completing the Claude 5.5 series. Additionally, Anthropic halved Sonnet 5.5’s cache-read price to $0.10 per million tokens and introduced monthly API credits for Max and Team subscribers.
The aggressive pricing strategy signals a shift toward commoditizing high-volume, low-complexity AI tasks. By matching OpenAI’s GPT-6 Luna rates and undercutting its own previous generation by 90%, Anthropic aims to capture market share in cost-sensitive applications such as customer support automation and document summarization. The introduction of adjustable effort settings further allows developers to optimize the trade-off between latency, cost, and performance, catering to diverse operational needs without requiring multiple model deployments.
While Haiku 5.5 demonstrates superior performance over direct competitors in agentic benchmarks like OSWorld and Terminal-Bench, the gap between it and the flagship Sonnet 5.5 remains significant in complex coding scenarios. This tiered capability structure reinforces a bifurcated market where smaller models handle routine infrastructure tasks, while larger models retain dominance in high-stakes reasoning. Enterprises must now evaluate whether the marginal cost savings justify potential accuracy risks in critical workflows, particularly given the source’s note on occasional logical errors despite rapid response times.


