Research investigation R0923 / benchmark analysis

Which Model Earned Opus 5.5’s Agent-Benchmark Score?

The published scores describe different evaluation configurations. Four selected Opus 5.5 rows permitted production-safeguard fallback to older Claude models; Zapier’s AutomationBench run did not. Permission to fall back does not show that fallback occurred.

Current public editionv1Sep 23, 2026
Verified observations
11

11 measured fields

Supported claims
9

9 material findings

Cited sources
8

8 primary or authoritative

Research score
88

Automated topic and evidence score

Interactive figureWhich Model Earned Opus 5.5’s Agent-Benchmark Score?
CSV JSON
Data status31.4 verified records across 1 period

Snapshot only. There is not enough history to claim a trend yet.

Verified observationHover or focus any mark for exact valuesLast updated Sep 23, 2026

Version ledger

Frozen public editions

Each edition preserves the records, method, sources, and downloads available at publication time.

  1. v1 / latestSep 23, 202611 records / 8 sources

    Initial public snapshot with 11 records and 8 cited sources.

Coverage note

Synthesis of supplied evidence snapshots retrieved September 23, 2026, at 15:34:22 UTC. That supplied snapshot time is later than the collection context’s stated 15:31:06 UTC currentTime; recovering it is not a new collection or experiment. The findings concern only the five selected rows.

Dataset ID
spd:which-model-earned-opus-5-5-s-agent-benchmark-score-d7a939a8
Stable URL
/research/which-model-earned-opus-5-5-s-agent-benchmark-score-d7a939a8
Version
v1
Coverage
2026-09-23
Records
11
Fields
7
Updated

Read the data

The records behind the figure

CSV JSON
Which Model Earned Opus 5.5’s Agent-Benchmark Score? data records
EntityMetricValueUnitObservedSourceTransform
Claude Fable 5.1 with Opus 5 fallbackAutomationBench 1.0.6 official private-set strict pass rate; max effort; fallback completions included31.4%percent2026-09-23https://zapier.com/benchmarks
Claude Opus 5.5AutomationBench 1.0.6 official private-set strict pass rate; max effort; no fallback; interventions fail40.0%percent2026-09-23https://zapier.com/benchmarks
Claude Opus 5.5CursorBench 4.0 coding-discussion score; medium effort52.5%percent2026-09-23https://www.anthropic.com/claude-opus-5-5
Claude Opus 5.5CursorBench 4.0 launch-table score; max-effort default; production safeguards and fallback permitted57.8%percent2026-09-23https://www.anthropic.com/claude-opus-5-5
Claude Opus 5.5FrontierCode v1.1 Main coding-discussion score; medium effort54.6%percent2026-09-23https://www.anthropic.com/claude-opus-5-5
Claude Opus 5.5FrontierCode v1.1 Main launch-table score; max-effort default; production safeguards and fallback permitted54.4%percent2026-09-23https://www.anthropic.com/claude-opus-5-5
Claude Fable 5.1 with Opus 5 fallbackTasks Zapier says Opus 5 handled in the Fable fallback configuration260 of 657 taskstasks2026-09-23https://zapier.com/benchmarks
GPT-6 AstraTerminal-Bench 4.0 OpenAI-reported comparator; high effort57.9%percent2026-09-23https://www.anthropic.com/claude-opus-5-5
Claude Opus 5.5Terminal-Bench 4.0 published resolution rate; xhigh effort; production safeguards and fallback permitted66.4%percent2026-09-23https://www.anthropic.com/claude-opus-5-5
Claude Opus 5Terminal-Bench-Science 0.1 benchmark-owner resolution rate; Claude Code; three trials per task30.0%percenthttps://www.tbench.ai/news/terminal-bench-science-0-1
Claude Opus 5.5Terminal-Bench-Science 0.1 published resolution rate; max-effort default; production safeguards and fallback permitted58.7%percent2026-09-23https://www.anthropic.com/claude-opus-5-5

Measurement technique

How to read this report

  1. 01Plan an evidence_matrix for the five selected rows: Terminal-Bench 4.0, FrontierCode v1.1 Main, CursorBench 4.0, Terminal-Bench-Science 0.1, and AutomationBench. Do not extend conclusions to Anthropic’s other benchmarks.
  2. 02For each reported score, record model or fallback configuration, safeguard rule, effort, benchmark version, harness, trials, evaluator, reporting party, and available task-level artifacts; mark undocumented fields unknown.
  3. 03Keep Anthropic’s launch-table scores distinct from its medium-effort coding-discussion scores, and compare benchmark-owner protocols without treating unlike grading methods as interchangeable.

Sources

Evidence

6 publishers supporting 11 records. Expand a publisher to inspect its cited pages.

Next report / 01AI Model Economics Index All research reports
YOUR READING SPACE

Notifications