Review Finds Only One of Five Opus 5.5 Benchmark Scores Rules Out Model Fallback

Zapier counted safeguard interventions as failures in its Opus 5.5 run. Four other rows permitted an older model to finish sensitive tasks, but their published scores do not reveal whether that happened.

By 5 min read
Original researchWhich Model Earned Opus 5.5’s Agent-Benchmark Score?

The published scores describe different evaluation configurations. Four selected Opus 5.5 rows permitted production-safeguard fallback to older Claude models; Zapier’s AutomationBench run did not. Permission to fall back does not show that fallback occurred.

Explore the full research
Review Finds Only One of Five Opus 5.5 Benchmark Scores Rules Out Model Fallback
Superpower DailyOriginal research
Review Finds Only One of Five Opus 5.5 Benchmark Scores Rules Out Model Fallback

Listen to this story

The audio brief

About 1:26
0:001:26
Read transcript
Zapier’s AutomationBench is the only one of five reviewed Opus 5.5 benchmark results that clearly rules out an older model stepping in. It scored Opus 5.5 at 40 percent, counting any safeguard intervention as a failure. That makes this result unusually clear about who did the work, but not a clean measure of what a safeguarded system might complete. The other four benchmarks allowed handoffs for certain sensitive tasks. Anthropic’s rules say Opus 4.8 could take cybersecurity work, while Opus 5 could take biology and frontier-model-development tasks. The published scores don’t say whether those handoffs actually happened, or how often. The comparison inside AutomationBench shows why the distinction matters. Zapier lists Claude Fable 5.1 at 31.4 percent, but that run included Opus 5 fallback. Zapier says Opus 5 handled 260 of 657 tasks. So the two percentages come from the same test, but they do not represent the same model-identity rule. This review covers coding, science, and workflow tests—not all of Anthropic’s benchmarks. The broader scores also vary with effort settings and reporting parties, so they aren’t all directly comparable. The unresolved question is narrow: did older models contribute to the other four results, and by how much? Task-level handoff records, followed by matched runs with fallback disabled, would answer it.

Story brief

3 key points

A review of Anthropic’s September 22 Opus 5.5 launch benchmarks finds that the reported scores do not all establish the same model identity: four selected coding and science tests allow older-model intervention, but public records do not show whether handoffs occurred. Zapier’s AutomationBench is the sole no-fallback case, scoring Opus 5.5 at 40.0% while counting safeguards as failures. Buyers should treat the other...

  1. 01

    AutomationBench lists Claude Fable 5.1 at 31.4% with Opus 5 fallback; Opus 5 handled 260 of 657 tasks, making it a different model-identity comparison.

  2. 02

    Anthropic’s fallback rules allow Opus 4.8 on cybersecurity tasks and Opus 5 on biology and frontier-model-development tasks; intervention counts remain unpublished.

  3. 03

    Terminal-Bench 4.0 compares Opus 5.5 at 66.4% with xhigh effort against GPT-6 Astra at 57.9% with high effort, reported by different companies.

A benchmark can carry one model’s name without showing which model finished every task. In a review of five Claude Opus 5.5 coding and workflow results, only Zapier’s AutomationBench explicitly barred fallback to an older model. Its 40.0% score counted safeguard interventions as failures. The other four rows permitted older models to take over specified sensitive work, but the published results do not show whether any handoff occurred.

That boundary affects what a score can describe. A fallback-permitted result may measure a safeguarded system rather than uninterrupted work by Opus 5.5. A no-fallback test makes model identity clearer, but counts a safeguard-stopped task as a failure. Neither rule alone predicts success on a buyer’s own work.

Five rows under one test of the labels

This review examines five rows in Anthropic’s September 22 launch table: Terminal-Bench 4.0, FrontierCode v1.1 Main, CursorBench 4.0, Terminal-Bench-Science 0.1 and AutomationBench. Using public documents available through September 23, it compares fallback rules, effort settings, graders and reporting parties. It does not classify Anthropic’s other benchmarks or claim to have rerun any test.

Anthropic’s table footnote says Opus 5.5 was evaluated with production safeguards. When they intervened, Opus 4.8 completed cybersecurity tasks; Opus 5 completed biology and frontier-model-development tasks. That rule applies to the four selected coding and science rows. Anthropic specifies an exception for AutomationBench: no fallback model, and an intervention meant failure. A permitted route is not evidence that a task actually took it.

Zapier’s clean boundary, uneven comparison

Zapier reports that max-effort Opus 5.5 completed 40.0% of AutomationBench’s official tasks. An agent works across simulated business apps, and fixed checks inspect the state it leaves behind. A task passes only when every scored requirement is met. The number measures strict workflow completion, not whether the agent wrote a convincing reply.

Anthropic’s table also lists Claude Fable 5.1 at 31.4%. Zapier’s leaderboard gives that entry a fuller name: Fable 5.1 with Opus 5 fallback. Zapier says Opus 5 handled 260 of 657 tasks after Fable’s safety classifier refused a step, and the score includes those fallback completions. The two percentages come from the same benchmark, but not the same model-identity rule.

Anthropic says counting interventions as failures lowered Opus 5.5’s score below what it would achieve in practice. That is a company counterfactual, not a measured fallback-enabled result presented here. Zapier’s released tasks cannot reproduce the official leaderboard figure exactly, either: the leaderboard uses a separate, harder private set.

World map showing the geographic density of Terminal-Bench-Science proposal and implementation authors using public profile locations and affiliations
376 contributors across 22 countries from proposals, reviews, or pull requests. Source: tbench.ai.

What the other four scores leave open

For the four remaining rows, the safeguard footnote identifies the possible handoff, not its frequency. The checked public materials give no task-level record dividing successful work between Opus 5.5 and an older model. They also provide no matched Opus 5.5 runs with fallback switched off. The results therefore support neither a claim that fallback boosted those scores nor a claim that Opus 5.5 completed every task alone.

The effort setting behind the percentage

Model identity is only one comparison setting. On Terminal-Bench 4.0, Anthropic puts Opus 5.5 at 66.4% using xhigh effort beside GPT-6 Astra at 57.9% using high effort. The Astra number was reported by OpenAI. These are published scores with different effort levels and reporting parties, not a matched-effort run by one evaluator. Anthropic’s Opus 5 reproduction scored 52.3%, close to the 51.8% public Claude Code leaderboard result it cites.

Terminal-Bench-Science 0.1 supplies another baseline check. Anthropic reports 58.7% for Opus 5.5 and 29.0% for its Opus 5 reproduction. The benchmark owner reports Opus 5 at 30.0%, with three independent Claude Code trials on each of 70 tasks. Those close Opus 5 figures put Anthropic’s reproduction in context, but do not reveal any Opus 5.5 handoffs. Anthropic attributes its table’s 64.6% Astra figure to OpenAI.

Even Anthropic’s own presentation gives different Opus 5.5 scores at different effort levels. FrontierCode v1.1 Main shows 54.4% in the max-effort launch table and 54.6% at medium effort in the coding discussion. CursorBench 4.0 shows 57.8% in the table and 52.5% at medium effort in that discussion. A reader should not silently substitute a discussion score for a headline score, or assume that higher effort always produces the higher reported number.

The next disclosure that would settle the question

The tests grade different kinds of work. Zapier checks the final state of a simulated business. Cursor’s published account of an earlier benchmark version describes agentic graders for ambiguous coding tasks, while Anthropic says FrontierCode assesses whether code changes would be merged. These methods add context to each percentage; they do not provide a common scale for ranking all five rows.

What to watch is narrower than another headline score: task-level intervention and handoff records for the four fallback-permitted rows, then matched results with fallback disabled. Those would show whether older models contributed to the reported outcomes and by how much. Until then, the AutomationBench result has a clear no-fallback boundary; the other four have a stated route, not an observed count.

Editorial analysis

Our Read

Our reading is that a benchmark score needs a model-identity rule alongside its model name. Zapier’s AutomationBench entry makes the distinction concrete: Opus 5.5 could not fall back, while the Fable 5.1 comparator includes work handed to Opus 5. Neither number is invalid, but the labels alone conceal a consequential difference. For the four other rows, the question is still open because permission to hand off a task does not show that a handoff occurred. Task-level intervention records, followed by matched runs with and without fallback, would show whether the safeguard route materially changes any of those scores.

Citation desk / original work

Cite this

Permanent attributionView citation
Finding 01

AutomationBench is the explicit exception: Zapier ran the Opus 5.5 evaluation without fallback models and counted safeguard interventions as failures. Its Opus 5.5 max-effort score is 40.0%, graded by strict, deterministic checks of the final simulated business state.

/posts/review-finds-only-one-of-five-opus-5-5-benchmark-scores-rules-out-model-fallback#finding-claim-03
Finding 02

Anthropic's general rule for its Opus 5.5 benchmarks enabled production safeguards: upon intervention, cybersecurity work could pass to Opus 4.8, and biology or frontier-LLM-development work to Opus 5. Thus the four selected coding/science rows permitted fallback; this does not establish that fallback occurred in every row or successful task.

/posts/review-finds-only-one-of-five-opus-5-5-benchmark-scores-rules-out-model-fallback#finding-claim-02
Finding 03

The Terminal-Bench 4.0 headline compares Opus 5.5 at xhigh effort (66.4%) with an OpenAI-reported GPT-6 Astra result at high effort (57.9%). Anthropic reports its own Opus 5 result as 52.3%, versus 51.8% on the five-trial-per-task public Claude Code leaderboard; the GPT figures were reported by OpenAI rather than generated in that Anthropic comparison.

/posts/review-finds-only-one-of-five-opus-5-5-benchmark-scores-rules-out-model-fallback#finding-claim-04

Sources

  1. zapier.comAutomationBench: AI Agent Benchmarks | Zapier
  2. github.comAutomationBench/README.md at main · zapier/AutomationBench
  3. tbench.aiwww.tbench.ai
  4. cursor.comHow we compare model quality in Cursor · Cursor
  5. anthropic.comIntroducing Claude Opus 5.5

Loading discussion...

YOUR READING SPACE

Notifications