
Anthropic released Claude Sonnet 5.5 on September 28. The company says the model responds faster and uses less computation for typical work, while keeping the same list price as Sonnet 5. For users, the useful question is whether a specific task can be completed more reliably and with less waiting, rather than whether the model selector displays a new version number.
Key takeaways
- Sonnet 5.5 still costs $2 per million input tokens and $10 per million output tokens through the API.
- Anthropic reports output that is more than 30 percent faster and task costs up to 30 percent lower in its own tests.
- An independent CodeRabbit evaluation found more bugs in code reviews, while also showing a meaningful gap to Opus 5.5.
- Teams migrating existing API applications need to check changes to thinking, tool calls, and streaming.
What changes in everyday work
Anthropic positions Sonnet 5.5 for well-defined everyday tasks, bug fixes, and the creation of documents, slides, and spreadsheets. The company says Opus 5.5 remains stronger when a problem demands sustained judgment over an open-ended process. That distinction is more useful than a blanket model ranking. A small script, a first code review, or a structured document calls for a different kind of work from designing the architecture of a large system.
The model is available through Claude and the Claude API. Anthropic’s developer documentation lists a one-million-token context window and a maximum output of 128,000 tokens. Those figures describe technical limits, not a promise that every answer based on a very long input will be accurate. Anyone processing a long case file or a large repository should check whether the model found, cited, and weighed the relevant passages correctly. A large context window does not replace careful selection of source material.
A useful first trial starts with a familiar task whose result can be checked, such as a small repository change with tests or a summary that can be traced back to original documents. Users can give the same prompt to their current model and compare the result, elapsed time, follow-up work, and corrections required. Our look at the Claude Pro plan can help with the separate question of whether app usage fits a particular workflow. API charges are a different cost category.
How an unchanged token price can reduce the bill
Anthropic charges the same $2 and $10 per million input and output tokens for Sonnet 5.5 as for Sonnet 5. Cache reads cost 20 cents per million tokens, according to the company. If a model completes the same work in fewer steps and with shorter responses, the cost per task can fall even when the rate card stays fixed. Anthropic says its own tests found savings of up to 30 percent per task and output generation more than 30 percent faster. Those results will not apply to every workload.
CodeRabbit compared the two Sonnet versions in its review pipeline. Among 13 difficult cases with known bugs, Sonnet 5.5 identified six problems through actionable comments, while Sonnet 5 identified four. Across 44 open-source pull requests, CodeRabbit measured an average review time of 6 minutes and 33 seconds rather than 13 minutes and 31 seconds. At list prices, the Claude model calls alone cost about 46 cents instead of $1.16 per review. That is a result for this particular pipeline, not a universal customer discount.
The evaluation states its own limits. Thirteen hard cases are a small sample, and scoring for bug coverage on the larger set was still pending when the article appeared. Sonnet 5.5 also missed two bugs that Sonnet 5 caught. Opus 5.5 did better on the same hard cases. Teams using automated reviews should therefore look beyond comment counts and check which real problems the model finds and which ones it misses, especially for high-impact changes.
How to read benchmarks and plan a migration
Anthropic reports a 70.6 percent score for Sonnet 5.5 on the agentic coding evaluation Terminal-Bench 4.0, compared with 10.3 percent for Sonnet 5. The gap is striking, but it comes from a defined test setup. It does not establish the same improvement in every project or general superiority over competing models. Anyone using the scores to make a purchasing decision needs to examine the test environment, effort setting, and actual cost per successful task together.
A model migration also involves more than changing a name in an API request. Anthropic documents changes to adaptive thinking, forced tool calls, and the way text between tool calls appears in a stream. An application that displays those events directly to users may behave differently after the switch. Testing the complete prompt and tool chain is therefore part of the migration, particularly when responses are streamed live or an agent uses several tools.
What matters next
Sonnet 5.5 looks like a useful option for frequent, well-defined work: the same token prices, faster output according to Anthropic, and lower cost per completed review in one independent pipeline. The decisive measure is still the result on a team’s own tasks. A sound pilot uses real examples with known outcomes, measures total costs including rework, and records failures. That shows whether the efficiency gains carry over to daily work without treating a benchmark as a guarantee.

