Research investigation R0923 / benchmark analysis
Which Model Earned Opus 5.5’s Agent-Benchmark Score?
The published scores describe different evaluation configurations. Four selected Opus 5.5 rows permitted production-safeguard fallback to older Claude models; Zapier’s AutomationBench run did not. Permission to fall back does not show that fallback occurred.
Snapshot only. There is not enough history to claim a trend yet.
Version ledger
Frozen public editions
Each edition preserves the records, method, sources, and downloads available at publication time.
Synthesis of supplied evidence snapshots retrieved September 23, 2026, at 15:34:22 UTC. That supplied snapshot time is later than the collection context’s stated 15:31:06 UTC currentTime; recovering it is not a new collection or experiment. The findings concern only the five selected rows.
- Dataset ID
- spd:which-model-earned-opus-5-5-s-agent-benchmark-score-d7a939a8
- Stable URL
- /research/which-model-earned-opus-5-5-s-agent-benchmark-score-d7a939a8
- Version
- v1
- Coverage
- 2026-09-23
- Records
- 11
- Fields
- 7
- Updated
Measurement technique
How to read this report
- 01Plan an evidence_matrix for the five selected rows: Terminal-Bench 4.0, FrontierCode v1.1 Main, CursorBench 4.0, Terminal-Bench-Science 0.1, and AutomationBench. Do not extend conclusions to Anthropic’s other benchmarks.
- 02For each reported score, record model or fallback configuration, safeguard rule, effort, benchmark version, harness, trials, evaluator, reporting party, and available task-level artifacts; mark undocumented fields unknown.
- 03Keep Anthropic’s launch-table scores distinct from its medium-effort coding-discussion scores, and compare benchmark-owner protocols without treating unlike grading methods as interchangeable.
Sources
Evidence
6 publishers supporting 11 records. Expand a publisher to inspect its cited pages.