Anthropic launches Claude Haiku 5.5 at $0.10 per million tokens, with effort controls and a 100k cliff
The cheapest Haiku yet targets high-volume and subagent pipelines, but the new tokeniser and the tiered pricing cliff over 100k input tokens decide what you actually save.
What happened
Summary of reporting by The DecoderAnthropic released Claude Haiku 5.5 on 7 October 2026, its smallest and fastest model, targeted at high-volume workloads: classification, extraction, summarisation, compaction, and narrow subagent tasks alongside Sonnet 5.5 and Opus 5.5. It is the first Haiku with adaptive thinking and an adjustable effort setting, plus a 1M token context window and up to 128k output tokens (300k on the Batches API with the output-300k-2026-03-24 beta header).
Pricing uses prompt-length tiers. Requests up to 100,000 input tokens cost $0.10 per million input and $0.50 per million output, compared to $1 and $5 on Haiku 4.5. Above 100k, rates rise to $0.50 and $2.50. Cache reads drop to $0.01 per million. Anthropic estimates average running-cost reductions at around 75%, partly offset because the new tokeniser emits roughly 30% more tokens for identical text. Alongside the launch, Sonnet 5.5 cache reads fall from $0.20 to $0.10 per million, worth roughly 20% on typical agentic runs, and Max and Team subscribers receive monthly API credits of $100, $200 and up to $500 pooled.
Reported capability gains are substantial: 72.4% on the OSWorld 2.1 offline subset against 15.7% for Haiku 4.5, 39.2% on Terminal-Bench 4.0 against 0%, and 1,620 on GDPval-AA v2.1 against 735. Early users named by Anthropic include AlphaSense, HubSpot, Box, Asana, and Cognition for Devin Fusion. The model is available as claude-haiku-5-5 on the Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry. Reporting: The Decoder (https://the-decoder.com/claude-haiku-5-5-arrives-with-massive-price-cuts-proving-the-ai-pricing-arms-race-is-far-from-over/).
The Azrty take
The cut is real, but the saving goes to teams that re-architect around cheap subagents and keep prompts under 100k tokens.
For GCC organisations running AI at volume, this shifts the economics of repetitive enterprise workloads: Arabic and English support triage, KYC and invoice extraction, document compaction, retrieval pre-filters, and the fan-out layer of multi-agent architectures. At $0.10 and $0.50 per million tokens, a Haiku-class call is cheap enough to run multiple times per document while still costing less than a single Sonnet pass. The Decoder notes the launch pricing matches OpenAI's GPT-6 Luna and frames it as a small-model price war. For enterprise buyers, the practical consequence is that cost per completed task, rather than nominal cost per model, is now the primary metric.
The capability jump makes this price cut practical rather than cosmetic. Haiku 5.5 scores 72.4% on the OSWorld 2.1 offline subset compared to 15.7% for Haiku 4.5, 1,620 on GDPval-AA v2.1 against 735, and 39.2% on Terminal-Bench 4.0 against 0%. Early deployments cited in Anthropic's announcement illustrate the intended deployment pattern: AlphaSense runs its Ask in Document feature at roughly 8 million calls weekly and measured 0.84 against 0.76 for Haiku 4.5 across 400 queries; HubSpot reports 92.8% on its CRM evaluation suite; Cognition placed Haiku 5.5 in the sidekick slot of Devin Fusion alongside Opus 5.5 as lead, reporting a FrontierCode score of 66.2. The architectural boundary remains clear: Sonnet 5.5 still scores 70.6% on Terminal-Bench 4.0, and Anthropic positions Haiku specifically for compaction, summarisation and subagent duties rather than complex autonomous coding.
Three operational traps determine whether teams realise the headline savings. First, token density: the Anthropic model documentation states the new tokeniser produces roughly 30% more tokens for the same English text compared to Haiku 4.5, meaning 1M tokens equates to about 555k words versus roughly 750k previously. Existing budget models carry an immediate 30% error before applying unit prices. Second, the 100k threshold: the $0.10 and $0.50 tier covers around 90% of historic Haiku requests, but any request exceeding 100k input tokens is billed at five times the baseline across input, output and cache under Anthropic's pricing documentation. Prompts hovering near 110k represent an expensive configuration defect. Third, data residency pricing: regional endpoints on Bedrock and Google Cloud Vertex carry a 10% premium over global pools, and pinning inference_geo to us on the Claude API applies a 1.1x multiplier across token classes, which affects UAE regulated deployments.
Our recommended approach treats this as an intelligent routing problem rather than a direct model swap. Place a gateway such as FastLLM Proxy in front of providers, configuring per-route model pinning, hard budgets and token limits. Triage, parsing and subagent extraction can land on claude-haiku-5-5, while primary agent loops remain on Sonnet 5.5 or Opus 5.5. Enable prompt caching on every system prompt and static context prefix, taking advantage of $0.01 per million cache reads. Set effort explicitly per route (the API default is medium) following the official effort parameter guide and run an evaluation sweep on internal traffic, particularly Arabic datasets, before migration. A representative subagent configuration using the Anthropic Python SDK looks like this:
import anthropic
client = anthropic.Anthropic()
Subagent extraction route: pinned model, low effort, cached system prompt
resp = client.messages.create( model="claude-haiku-5-5", # Bedrock: anthropic.claude-haiku-5-5 max_tokens=2048, output_config={"effort": "low"}, # default is medium; sweep low and medium on your evals system=[{ "type": "text", "text": EXTRACTION_SYSTEM_PROMPT, "cache_control": {"type": "ephemeral"}, # cache reads at $0.01/MTok }], messages=[{"role": "user", "content": doc_text[:90_000]}], # keep strictly under the 100k tier )
What to do now
- Pin claude-haiku-5-5 (or anthropic.claude-haiku-5-5 on Bedrock) for classification, extraction, compaction and subagent tasks, keeping claude-sonnet-5-5 on primary coding loops where Terminal-Bench 4.0 performance remains 70.6% versus 39.2%.
- Re-tokenise your top 10 prompts and run your evaluation suite across low and medium effort settings before migrating traffic. The new tokeniser produces roughly 30% more tokens, meaning legacy budget forecasts overstate savings.
- Add cache_control breakpoints to system prompts and document prefixes, and enable 1-hour prompt caching on Sonnet 5.5 agent sessions following the cache read price reduction from $0.20 to $0.10 per MTok.
- Enforce an automated 100k input-token ceiling in your gateway to prevent the 5x price escalation on larger prompts. Account for the 10% regional endpoint premium or the 1.1x inference_geo multiplier if UAE residency rules apply.
