Anthropic ships Claude Sonnet 5.5: near-Opus scores at half the price, but a breaking API migration
Same list price as Sonnet 5 ($2/$10 per million tokens), 30%+ faster output, and up to 30% lower cost per task. The model ID swap is the easy part; thinking defaults, tool_choice and token accounting all changed.
What happened
Summary of reporting by AnthropicAnthropic released Claude Sonnet 5.5 on 28 September 2026. It is the second model in the Claude 5.5 family after Opus 5.5. List pricing matches Sonnet 5 at $2 per million input tokens, $10 per million output tokens, and $0.20 per million cache reads. Anthropic reports the model writes over 30% faster. It also needs fewer tokens per task, cutting effective cost per task by up to 30%. Claude Haiku 5.5 will follow in the coming weeks.
The headline leap comes in agentic coding. On Terminal-Bench 4.0, Sonnet 5.5 reaches 70.6%, up from Sonnet 5's 10.3%. It sits near Opus 5.5 on CursorBench 4.0 (55.5% vs 57.8%) and GDPval-AA v2.1 (1,844 vs 1,846). Crucially, it runs much cheaper per task at low and medium effort. Anthropic notes Opus 5.5 still leads on open-ended work needing sustained judgment. The lab also warns that Sonnet 5.5 scores worse at Max effort than Xhigh on FrontierCode. At Max effort, multi-agent reviews time out or edit out of scope. The model is live on the Claude Platform as claude-sonnet-5-5, and on AWS Bedrock, Google Cloud and Microsoft Azure.
Two operational changes matter to engineering teams. First, Sonnet 5.5 adds cybersecurity safeguards and anti-distillation classifiers. High-risk cyber requests drop back to Sonnet 5 automatically. Preserved thinking is now bound to the originating account and conversation. Second, Anthropic issued new prompting guidance for Opus 5.5. Developers should retest effort settings instead of reusing Opus 5 configs. They should also drop lines like 'think carefully', because the model now allocates its own reasoning.
Reporting: TechCrunch (https://techcrunch.com/2026/09/28/anthropic-releases-sonnet-5-5-which-it-calls-a-significantly-cheaper-faster-work-partner/), The Decoder (https://the-decoder.com/anthropics-claude-sonnet-5-5-nearly-matches-opus-5-5-on-benchmarks-while-costing-up-to-30-percent-less-per-task/) and Search Engine Journal on the Opus 5.5 prompting guide (https://www.searchenginejournal.com/anthropic-claude-opus-5-5-prompting-guidance/591278/).
The Azrty take
The unit economics of agent workloads just moved decisively. The migration is where most teams will lose the saving.
Look at cost per completed task, not cost per token. That is the figure that governs your budget. Sonnet 5.5 keeps Sonnet 5's list price: $2 per million input tokens and $10 per million output tokens. That is half of Opus 5.5 ($4/$20). Yet Anthropic reports up to 30% lower cost per task. The model uses fewer tokens overall and batches tool calls more aggressively. Early deployment data confirms this trend. Slack saw roughly 14% fewer output tokens on Slackbot evals with no prompt edits. Zendesk resolved support tickets 20% faster. Base44 matched Opus 5 quality across 118 app builds in 3.6 iterations per build, compared to 7.7 for Opus 5. GCC organisations have many pilots stalled by Opus pricing. For those teams, this release clears the hurdle into production.
The benchmark gains are real, but check the footnotes. Terminal-Bench 4.0 (https://github.com/harbor-framework/terminal-bench) jumps from 10.3% to 70.6%. That slightly beats Opus 5.5's 66.4% at Xhigh effort. GDPval-AA v2.1 rates Sonnet 5.5 at 1,844 against Opus 5.5's 1,846 and GPT-6 Sol's 1,487. Do not treat Max effort as your default, though. On FrontierCode 1.1, Sonnet 5.5 drops from 52.1% at Xhigh to 46.2% at Max effort. At Max, it spawns multi-agent code reviews that time out or make out-of-scope edits. In addition, the GDPval-AA and AA-Briefcase numbers came from a pre-release run with a structured-output bug that is now patched. The 'near-Opus' claim holds for well-scoped tasks. It fails for open-ended strategic judgment. Watch the Artificial Analysis charts (https://artificialanalysis.ai/models) as independent output-token data arrives.
Most teams will stumble by treating this release as a simple model ID swap. The official migration guide (https://platform.claude.com/docs/en/models/sonnet-5-5/migration-guide) contains several hard 400 errors. Setting thinking.type: "disabled" now breaks; between_tools is the new floor. Unspecified thinking runs by default. Forced tool choice (tool_choice set to any or tool) is gone. You must use auto with strict: true under JSON Schema rules (https://json-schema.org/specification). Custom temperature, top_p and top_k parameters now error out. Assistant prefill is forbidden. Thinking tokens also count against max_tokens and carry output billing. An old token limit sized for Sonnet 5 will truncate your output. Cache limits dropped from 1,024 tokens to 512. Images can consume up to 4,784 visual tokens. Crucially, thinking blocks are cryptographically tied to conversation prefixes (https://platform.claude.com/docs/en/build-with-claude/preserved-thinking). If your client trims message history or re-encodes images between turns, requests will fail. Cloud platforms lag slightly behind: Amazon Bedrock (https://docs.aws.amazon.com/bedrock/latest/userguide/model-cards-anthropic.html) lacks strict tool choice at launch, so client-side schema validation remains necessary.
Treat this upgrade as a routing and quality challenge, not a simple version bump. Funnel all model traffic through an OpenAI-compatible gateway. That keeps routing, team budgets and access controls in one place. We built our FastLLM Proxy to handle this across 80 hosted providers. Route by task type. Use Sonnet 5.5 at medium effort for routine documents and defined coding runs. Switch to high effort for thorny builds. Reserve Opus 5.5 for open-ended analysis. Gate coding output before any code hits a repository. Cheaper tokens must not erode code quality. We use Procoder for this guardrail: untested or unformatted code cannot merge. A migrated request should look like this:
response = client.messages.create(
model="claude-sonnet-5-5", # no date suffix; pin it explicitly
max_tokens=16000, # re-size: thinking counts against this
thinking={"type": "between_tools"}, # "disabled" now returns 400
output_config={"effort": "high"}, # levels recalibrated: re-sweep low/medium/high
messages=[{"role": "user", "content": "..."}],
)
Two regional realities stand out for Gulf enterprises. First, the unit economics reward volume. Teams in Dubai, Abu Dhabi and Riyadh running high-throughput processing will pocket the 30% savings immediately. Haiku 5.5 will widen that gap shortly. Second, prepare for the new safeguards. Sonnet 5.5 adds active anti-distillation classifiers and automatic fallbacks to Sonnet 5 on cyber requests. Security teams running offensive testing need clearance through Anthropic's Cyber Verification Program. Multi-model routers that switch models mid-chat must pass thinking blocks back untouched, or the context drops. The opportunity is substantial. Do not blow it with a lazy model-string swap.
What to do now
- Audit Messages API calls immediately. Replace `thinking.type: "disabled"` with `between_tools`. Remove thinking budgets and custom temperature, top_p or top_k values. Switch forced `tool_choice` to `auto` with `strict: true`.
- Pin `claude-sonnet-5-5` explicitly. Re-sweep effort at low, medium and high across real tasks before going live (the API defaults to high). Re-size `max_tokens` to absorb thinking output.
- Keep multi-turn conversations strictly append-only. Pass thinking blocks back without modification. Add the `thinking-binding-controls-2026-08-01` header with `prefix_mismatch_behavior: drop_block`, and track `input_transformations`.
- Centralise agent traffic behind an API gateway with per-team spend caps, like FastLLM Proxy across 80 providers. Enforce quality gates on coding agents with Procoder. Security teams must register with the Cyber Verification Program before running high-risk workloads.
