Anthropic Releases Sonnet 5.5, Claiming Lower Task Costs Without Cutting Token Prices
The company says fewer tokens and tool calls can cut the cost of finished work by up to 30%. The result will depend on the task and effort setting.
Loading page…
The company says fewer tokens and tool calls can cut the cost of finished work by up to 30%. The result will depend on the task and effort setting.
Listen to this story
Anthropic’s September 28 launch adds Claude Sonnet 5.5 as a lower-cost-per-completed-task option for defined work, while leaving per-token API pricing unchanged. The company attributes savings of up to 30% to fewer tokens and tool calls, and says generation is over 30% faster than Sonnet 5; neither claim guarantees lower costs for every customer. Early tests from Box, Zendesk and Lovable report workflow-specific gains. Buyers should benchmark completed tasks at their intended effort settings, which differ across An
Sonnet 5.5 scored 70.6% on Terminal-Bench 4.0, versus 10.3% for Sonnet 5 and 66.4% for Opus 5.5.
Box reported 2.4× faster runs and 12% fewer tokens; Lovable reported one-third fewer tool calls and half as many shell executions.
Higher-risk cyber requests may fall back to Sonnet 5; separate classifiers aim to prevent large-scale extraction of the model’s reasoning.
A faster Sonnet model is now Anthropic’s alternative to paying for its flagship on everyday work. Released September 28, Claude Sonnet 5.5 carries the same API rates as its predecessor, but Anthropic says it can cut the cost of a finished task by up to 30% through fewer tokens and calls to outside tools.
Opus 5.5 arrived the week before. Anthropic recommends that flagship for ambiguous, open-ended work and positions Sonnet 5.5 for more defined jobs, including debugging software, making documents and refining interfaces. It says Sonnet also generates output more than 30% faster than Sonnet 5. Faster output alone does not establish the cost of completing a job.
Anthropic reports a 70.6% score for Sonnet 5.5 on the Terminal-Bench 4.0 coding test, against 10.3% for Sonnet 5 and 66.4% for Opus 5.5 under the reported settings. On a test drawing tasks from real-world occupations, Sonnet scored 1844 against Opus’s 1846. Those company results do not predict performance on every customer job.
Early customer tests offer narrower examples. Box said the model was more accurate, ran 2.4 times faster and used 12% fewer total tokens than its predecessor. Zendesk said it processed support tickets 20% faster and made fewer incorrect decisions than Claude models it currently uses in production. Lovable said its coding evaluations needed about one-third fewer tool calls and half as many shell executions. Each result comes from that customer’s testing.
Anthropic says this is its first Sonnet model with cybersecurity safeguards modeled on those for its highest-capability systems. Routine development and vulnerability fixes should work normally, it says, but some higher-risk requests will fall back to Sonnet 5. The company plans to widen a verification program for approved defenders. Sonnet 5.5 also adds classifiers intended to block extraction of its reasoning through large-scale queries—protection against copying the model, rather than the cyber-request risk.
Anthropic says Sonnet 5.5 will be available through its services, Amazon Web Services, Google Cloud and Microsoft Azure, with zero-data-retention support. Developers can call it as claude-sonnet-5-5. Haiku 5.5, intended for high-volume and especially cost-sensitive work, is planned for the coming weeks.
Claude Code and Anthropic’s consumer apps default to Medium effort, while its developer platform defaults to High. Lower effort uses less time and fewer tokens; higher effort lets the model check and refine its work longer. That choice could change both the bill and the result, so the useful comparison is a completed task at the setting a customer would actually use.
Editorial analysis
Sonnet 5.5 puts a practical question behind Anthropic’s model tiers: how much checking does a task need before the cheaper option stops being cheaper? The company recommends Opus 5.5 for open-ended work, yet its coding-test result puts Sonnet ahead under the reported settings. Those facts can coexist because a benchmark score cannot price the retries or review a particular job requires. The next useful evidence would be completed-work comparisons on the same tasks at the effort settings customers actually use, including cases where a Sonnet result needs another attempt or a switch to Opus.
Loading discussion...
Join the conversation
Which choice would you want for your everyday tasks?
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.