Claude Sonnet 5.5 Migration Guide: Six API Changes to Test

Claude Sonnet 5.5 keeps Sonnet 5 pricing but changes thinking, tool choice, computer use, conversation replay, and response handling. Test these six issues before switching production traffic.
Claude AI symbol used for coverage of Anthropic models, APIs, and agent tools.
Claude AI symbol. Image source: Wikimedia Commons; public-domain/CC0 asset.

Anthropic released Claude Sonnet 5.5 on September 28 as a faster replacement for Claude Sonnet 5, but developers should not treat the upgrade as a one-line model-name change. Existing API clients can fail with HTTP 400 errors if they disable thinking, force a tool call, reuse the older computer-use tool, or replay a modified conversation containing preserved thinking blocks.

The new model is available through the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS. It keeps Sonnet 5’s list price of $2 per million input tokens and $10 per million output tokens, while Anthropic claims it produces output more than 30% faster and can cost up to 30% less per completed task by using fewer tokens and tool calls. Those are vendor-reported results, so production teams should verify them against their own prompts, tools, latency targets, and retry patterns.

For teams already using Sonnet 5, the safest upgrade path is a staged migration that tests six compatibility changes before shifting production traffic.

Claude Sonnet 5.5 pricing and limits

Item Claude Sonnet 5.5
Model ID claude-sonnet-5-5
Input $2 per million tokens
Output $10 per million tokens
Cache read $0.20 per million tokens
Five-minute cache write $2.50 per million tokens
One-hour cache write $4 per million tokens
Batch API 50% discount on input and output
Context window 1 million tokens
Maximum output 128,000 tokens; 300,000 in Message Batches with a beta header
Default API effort high

The token rates are unchanged from Sonnet 5. The claimed savings come from completing work with less generation and fewer actions, not from a lower meter price. That distinction matters for agent systems: a faster model can still cost more if a prompt change causes longer reasoning, repeated tool calls, or retries.

1. Replace disabled thinking with between_tools

Claude Sonnet 5 accepted thinking: {"type": "disabled"}. Sonnet 5.5 rejects that value. To suppress up-front thinking, use between_tools at low, medium, or high effort:

{
  "model": "claude-sonnet-5-5",
  "max_tokens": 4096,
  "thinking": {"type": "between_tools"},
  "output_config": {"effort": "medium"},
  "messages": [
    {"role": "user", "content": "Summarize this incident report."}
  ]
}

between_tools does not work at xhigh or max. Those levels require adaptive thinking. Manual thinking budgets are also rejected, so applications that previously set budget_tokens need to move to the effort control.

Thinking also affects capacity planning. The API counts thinking and visible text against max_tokens, and thinking tokens are billed as output. Recheck limits on responses that were already close to the ceiling.

2. Stop forcing individual tool calls

Sonnet 5.5 does not support tool_choice values of any or a named tool. Both return an invalid-request error. The supported choices are auto, which is the default, and none.

If an integration used forced tool choice to guarantee structured arguments, Anthropic recommends auto with strict tool use. If the application requires structured data rather than an actual action, structured outputs are the cleaner fit. Prompts can still state when a tool should be used, but teams should test the failure path rather than assuming the model will always call it.

3. Read response blocks by type

Applications that assume response.content[0].text can break because a response may begin with a thinking block. Tool loops should inspect every content block by its type, render text blocks, execute tool-use blocks, and pass thinking blocks back unchanged.

There is also a quieter interface failure to watch. Longer progress notes between tool calls now arrive as thinking blocks. With the default display setting, their text is omitted. An agent can therefore appear silent while it continues working. Products that show intermediate status should either request summarized thinking display where appropriate or use between_tools, which returns the progress text.

4. Keep conversation history append-only

Sonnet 5.5 binds preserved thinking blocks to the model, account, and preceding conversation. On newer accounts, replaying a thinking block after changing the system prompt, tool definitions, or an earlier message can produce a 400 error.

That behavior affects systems that edit chat history to inject fresh instructions, redact an earlier turn, or swap tools while retaining old reasoning. The straightforward pattern is append-only history: add a new system message or a supported mid-conversation tool change instead of rewriting prior content. If an application must modify earlier history, it should remove affected thinking blocks or use Anthropic’s binding-control option to drop mismatched blocks.

Model routing needs its own test. Sonnet 5.5 can read thinking blocks produced by Sonnet 5 and several older Claude models, but other models do not read Sonnet 5.5 thinking blocks. A fallback or mid-conversation model switch may succeed after silently dropping that context.

5. Update computer use integrations

On the Claude API and Google Cloud, Sonnet 5.5 rejects the older computer_20251124 tool. Integrations must use computer_toolset_20260801 and update their agent loop for member tool-use blocks, batch actions, and the toolset_name returned with results. Amazon Bedrock still accepts the older tool, making this a platform-specific migration rather than a universal change.

Browser-use integrations and clients already using the newer toolset do not need that conversion. Teams operating across more than one cloud should keep platform capability tests in deployment rather than sharing an unchecked tool declaration everywhere.

6. Recalibrate effort instead of copying the old setting

The effort levels in Sonnet 5.5 do not map directly to the amount of thinking used by Sonnet 5. Anthropic suggests starting at medium for well-specified agentic or latency-sensitive tasks and high for harder, longer work. The Claude API defaults to high, while Claude Code and Anthropic’s consumer apps default to medium.

Run an effort sweep on representative jobs and record task success, end-to-end latency, total input and output tokens, tool-call count, failure rate, and human correction time. Per-token pricing alone will not show whether the upgrade pays off.

How Sonnet 5.5 compares with Opus 5.5

Anthropic positions Sonnet for well-defined everyday work and Opus for ambiguous tasks that require sustained judgment. Its published results put Sonnet 5.5 close to Opus 5.5 on several tests: 1,844 versus 1,846 on GDPval-AA, 80.1% versus 81.8% on the partial OSWorld 2.1 computer-use evaluation, and 55.5% versus 57.8% on CursorBench 4.0.

The eye-catching exception is Terminal-Bench 4.0, where Anthropic reports 70.6% for Sonnet 5.5 and 66.4% for Opus 5.5 under the disclosed settings. That does not establish that Sonnet is generally more capable. Effort levels and test harnesses can change both cost and scores, and Anthropic still recommends Opus for difficult, open-ended work.

The practical routing decision is workload-specific. At list price, Opus 5.5 costs twice as much per input and output token. Sonnet can be the default for bounded coding, support, extraction, and document tasks, with Opus reserved for cases that fail a quality threshold or require broader judgment.

A production migration checklist

  • Change the model ID in a staging environment.
  • Search the codebase for thinking, budget_tokens, tool_choice, content[0], and computer_20251124.
  • Test tool loops with thinking blocks before, between, and after tool calls.
  • Test edited histories, model fallbacks, and cross-account conversation transfer.
  • Compare low, medium, and high effort on the same evaluation set.
  • Measure completed-task cost, not only tokens per request.
  • Canary a small share of production traffic and retain a rollback path.
  • Monitor 400 errors, refusals, empty progress displays, tool retries, latency, and task completion.

Anthropic also introduced stronger cyber safeguards for Sonnet 5.5. High-risk cybersecurity requests may be refused or routed through a fallback, while routine software development and vulnerability remediation are intended to continue normally. Security products should test legitimate defensive workflows and handle stop_reason: "refusal" plus the returned policy category rather than treating every HTTP 200 response as a completed task.

The upgrade is attractive because it promises better capability at the same token rates, but its value will come from the entire agent loop. Teams that validate thinking, tools, history, effort, and fallbacks are more likely to see the claimed speed and cost gains without turning deployment day into an API debugging exercise.

Sources

Leave a Reply

Your email address will not be published. Required fields are marked *

Previous Post
OpenAI GPT-6.1 Sol model icon

GPT-6.1 Sol Costs One-Fifth of Astra, but Long Context Changes the Math

Next Post
Cloudflare illustration for its post-quantum certificate authority announcement

Cloudflare Plans a Post-Quantum Certificate Authority: What Changes for HTTPS

Related Posts