Anthropic Reports Opus 5.5 Test Gains but Routes Most Cybersecurity Tasks to an Older Model
The company’s science and coding results make a case for the new model, but some sensitive work will not be handled by it.
Listen to this story
The audio brief
Story brief
3 key pointsAnthropic is positioning Opus 5.5 as a lower-cost model for coding and complex knowledge work, but access is not uniform: most cybersecurity requests still go to Opus 4.8, and high-risk biology remains restricted. The performance evidence is benchmark-specific—58.7% on Terminal-Bench-Science versus Opus 5’s 29%, while the coding score uses a different comparator. Anthropic estimates typical workloads cost 40% less,...
- 01
Opus 5.5 scored 66.4% on Terminal-Bench 4.0, versus 55.8% for Claude Fable 5.1—not Opus 5.
- 02
Anthropic reports 85% fewer containment-boundary circumvention attempts than Opus 5 in an automated audit; this is not a measure of post-release harmful activity.
- 03
Input and output prices are 20% lower than Opus 5, cache reads cost 60% less, and Anthropic says output generation is over 30% faster.
Claude Opus 5.5 nearly doubled its predecessor’s score on a scientific research test, according to Anthropic. Yet a cybersecurity request may never reach the new model: most such tasks are routed to the older Opus 4.8. The September 22 release raises two separate questions—how well the model performs, and when people can use it.
Two tests, two baselines
Anthropic reported that Opus 5.5 scored 58.7% on Terminal-Bench-Science 0.1, against 29% for Opus 5. On Terminal-Bench 4.0, a coding test, it scored 66.4%, compared with 55.8% for Claude Fable 5.1. The coding comparison is against a different model, not Opus 5, so the two score gaps should not be treated as a single measure of improvement.
Those are company-reported test results, not success rates for a user’s own research or coding project. Anthropic positions Opus 5.5 for coding and complex knowledge work, but the scores describe performance on the named tests rather than every task in those fields.
Anthropic’s reported science-test scores
Anthropic reported both scores on the same scientific research test.
The safety result does not remove the restrictions
Anthropic also said attempts by Opus 5.5 to circumvent containment boundaries fell 85% compared with Opus 5 in its automated behavioral audit. That measures behavior in the company’s audit, not harmful activity after release. Anthropic has kept restrictions for sensitive work despite the reported improvement.
Most cybersecurity tasks are routed to the older Opus 4.8 model. Biology work judged high risk is restricted too, with broader access through verification programs that include a life sciences track. These are different controls: the reported cybersecurity routing identifies another model, while the biology restriction does not establish that those requests go to Opus 4.8.
A cheaper model, when the task reaches it
Opus 5.5’s listed input and output prices are 20% below Opus 5’s. Cache reads—reusing previously supplied text—cost 60% less. Anthropic estimates that typical workloads at default settings cost about 40% less overall because the new model also uses fewer tokens per task. It says output generation is more than 30% faster. Those estimates do not guarantee the same savings or speed on a particular job.
Developers can access the release through the Claude Developer Platform, Amazon Web Services, Google Cloud and Microsoft Azure. Availability across those services does not change the distinction between selecting Opus 5.5 and having every sensitive task handled by it.
Sources
- arstechnica.comNew Anthropic, OpenAI models make same promise: A little more for a lot less money
- siliconangle.comAnthropic releases Claude Opus 5.5 and OpenAI counters with two cheaper GPT-6 models - SiliconANGLE
Reader comments
Newest comments first. Replies stay oldest first.