Claude Sonnet 5.5 arrived on September 28, 2026, with a claim that sounds unusually attractive: better performance without a higher API price. Anthropic says the model generates output more than 30% faster than Sonnet 5 and can cost up to 30% less per task in its tests.
But the price of a model and the cost of getting useful work done are different things. Independent testing of Sonnet 5.5 at its maximum effort setting reached a different result: a higher cost per benchmark task than Sonnet 5. So is the new model cheaper? The answer depends on what you ask it to do, how much reasoning you allow and how you measure a completed task.
What Anthropic has announced
Sonnet 5.5 is the second model in Anthropic’s Claude 5.5 family, following Opus 5.5. Anthropic positions Sonnet as a faster option for clearly defined work, including fixing bugs and creating documents, presentations and spreadsheets. It describes Opus as better suited to complex, open-ended work that requires sustained judgment.
The new Sonnet model is available through Claude and its API. Anthropic also lists availability through Amazon Web Services, Google Cloud and Microsoft Azure. For developers, its Claude API model ID is claude-sonnet-5-5.
Its standard API rates remain the same as Sonnet 5: $2 per million input tokens and $10 per million output tokens. Anthropic says Sonnet 5.5 often needs fewer tokens to finish the same work, which is how a task could become cheaper even though the rates have not changed.
Why the price per token does not tell the whole story
An API bill depends partly on how many tokens a task consumes. A longer prompt, more reasoning, repeated tool calls or a lengthy answer can change the total. An agent that has to retry a failed step can cost more again.
Consider a simplified example. At Sonnet 5.5’s standard rates, a task using 100,000 input tokens and 20,000 output tokens would cost $0.40. If another run achieved the same result using 80,000 input tokens and 16,000 output tokens, it would cost $0.32. That is a hypothetical illustration, not a measured saving for Sonnet 5.5.
It also leaves out an important question: did both runs produce work of the same quality? A cheap answer that needs extensive correction may be more expensive in practice than a thorough answer that uses more tokens. For a business or developer, a useful comparison should account for successful completion, retries, review and time spent waiting.
Anthropic’s tests and an independent result point in different directions
Anthropic reports that Sonnet 5.5 costs up to 30% less per task than Sonnet 5 in its testing. It also says the new model generates output more than 30% faster. These are the company’s results; the “up to” figure should not be read as a saving on every request.
Artificial Analysis tested the model independently and found a different trade-off at maximum effort. Sonnet 5.5 scored 56 on its Intelligence Index, two points behind Opus 5.5 at maximum effort. But it also used an unusually large number of output tokens and cost $7.60 per task in that evaluation, approximately 50% more than Sonnet 5 under the comparison reported by Artificial Analysis.
These findings do not establish that either measurement is wrong. They examine different workloads and settings. At maximum effort, a model may spend substantially more tokens pursuing a stronger result. Artificial Analysis also reports considerably lower costs for Sonnet 5.5 at its lower effort settings. Its figures therefore show a range of possible outcomes, rather than a single price for every task.
The independent tests used a version deployed before the public release. Artificial Analysis notes that Anthropic subsequently fixed a bug that could affect some requests using structured outputs, and that relevant evaluations may be rerun. That limitation matters when interpreting close performance comparisons.
What changes for everyday users and developers?
For someone using Claude through its apps, the API’s price per million tokens is not a direct measure of their subscription cost. The practical questions are whether Sonnet 5.5 finishes a task faster, produces a better first draft and requires fewer corrections.
For teams building with the API, the effort setting deserves particular attention. A higher setting can give the model more room to work through a difficult problem, but it can also increase the number of tokens and the bill. Running every request at the highest setting may erase the savings expected from a more efficient model.
A useful trial would give Sonnet 5 and Sonnet 5.5 the same representative tasks, then record the quality of the result, total tokens, elapsed time, retries and human review required. Test straightforward requests separately from difficult ones. If Opus 5.5 is also an option, compare it on the complex tasks where extra judgment could justify its higher token rates.
The useful question is what a successful result costs
Sonnet 5.5 offers a clear improvement on paper: Anthropic has kept Sonnet’s API rates while reporting gains in speed and capability. Its claim of a lower cost per task may hold for many workflows. Independent testing shows that it should not be treated as a universal discount, especially when the model runs at maximum effort.
The most useful measure is the total cost of a result you can actually use. That includes the model’s tokens and, where relevant, the extra runs and human work needed to get the task finished.
