Anthropic’s new Claude Haiku 5.5 lists at a tenth of its predecessor’s price per token when a prompt stays within the 100,000-token line, and half the price above it. The company released it on 7 October. The cheaper small model is not the only price cut in the announcement, and it may not be the one that moves your invoice.
In its announcement post, Anthropic says Haiku 5.5 costs “around 75% less to run” than Haiku 4.5 on average, and it footnotes the claim. Read that as a blended estimate, not a rate card. The pricing table gives the actual list rates: $0.10 per million input tokens and $0.50 per million output tokens for shorter prompts, against $1.00 and $5.00 for Haiku 4.5. Anthropic says prompts under 100,000 tokens made up about 90% of requests to the old Haiku.
The company is not pitching Haiku 5.5 as a replacement for its larger models. It describes a worker for summaries, context compaction, database queries and classification, and a subagent that takes the small jobs while Opus 5.5 and Sonnet 5.5 handle coding. The post says outright that the bigger models are the ones to use for hard coding agents.
Every benchmark figure in the post is Anthropic’s own, taken from its own table, and none has been checked by a third party. By that yardstick the gap to Sonnet 5.5 is plain. On Terminal-Bench 4.0, a test of long command-line tasks, Haiku 5.5 scores 39.2% to Sonnet’s 70.6%. On OSWorld 2.1 (offline subset), where agents operate a computer, it scores 72.4% to 83.9%, up from 15.7% for Haiku 4.5.
The table also includes one outside model, GPT-6 Luna, placed between Haiku 4.5 and Sonnet 5.5. Haiku 5.5 beats it on the four tests cited here: 72.4% against 48.9% on OSWorld, 39.2% against 16.4% on Terminal-Bench, 46.4% against 42.4% on FrontierCode 1.1, and 46.4% against 29.1% on Chartography without tools. A vendor that names a single rival as its only outside yardstick shows whom it thinks small-model buyers are comparing it with. The table lists no Luna price, so the cost half of that contest is missing.
Sonnet 5.5 gets cheaper too. Cache reads, the discounted rate for context the model has already seen, drop from $0.20 to $0.10 per million tokens, effective immediately. Agents resend long context at nearly every step, so Anthropic puts the saving at about a fifth on typical agent jobs, which it states as “around 20% cheaper on most agentic work” for Sonnet 5.5. That figure is the company’s estimate, and it will shift with how cache-heavy your workload is.
Subscribers get money as well. Starting this week, Anthropic is rolling out a monthly API credit: $100 for Max 5x, $200 for Max 20x, and up to $500 for Team plans, pooled across members. The credits work on any model on the Claude Platform, which gives a flat-rate subscription a built-in budget for prototyping agents that call the API, Haiku 5.5 subagents included.
Smaller details sit further down. Haiku 5.5 is the first Haiku-class model with an adjustable effort setting, so you can trade cost against quality. Asana, quoted in the post, reports over 30% lower latency and up to 2.5x faster inference per agent turn than the model it uses today. Anthropic’s cybersecurity safeguards on the model block penetration testing. It is available on Amazon Web Services, Google Cloud and Microsoft Azure as claude-haiku-5-5.
Pull last month’s bill and split out cache reads before anything else. For an agent-heavy team on Sonnet, the 20% cut arrives without a code change, while Haiku 5.5 only pays off where you can carve out the small jobs.
Anthropic, in its own announcement post “Introducing Claude Haiku 5.5”, published 7 October 2026.