Research investigation R0904 / benchmark analysis

Astra’s Computer-Use Claims: How Much Is Model, and How Much Is Harness?

Astra may be highly capable, but the disclosed results do not consistently isolate model performance from the surrounding agent system.

Archived snapshotv2Sep 6, 2026
Verified observations
29

17 measured fields

Supported claims
9

9 material findings

Cited sources
11

11 primary or authoritative

Research score
83

Automated topic and evidence score

Interactive figureAstra’s Computer-Use Claims: How Much Is Model,...
CSV JSON
Data status59.3 verified records across 1 period

Snapshot only. There is not enough history to claim a trend yet.

Verified observationHover or focus any mark for exact valuesLast updated Sep 6, 2026

Version ledger

Frozen public editions

Each edition preserves the records, method, sources, and downloads available at publication time.

  1. v2 / latestSep 6, 202629 records / 11 sources

    Source-backed fields changed; the frozen snapshot contains 29 records.

  2. v1Sep 5, 202629 records / 11 sources

    Initial public snapshot with 29 records and 11 cited sources.

Coverage note

The ledger covers nine launch-table or independently reported Astra rows and excludes ARC-AGI-3. Field coverage is highest for evaluator or benchmark version (6 of 9) and task split (5 of 9), while state management and hard time/token budgets are disclosed for none of the nine rows.

Dataset ID
spd:astra-s-computer-use-claims-how-much-is-model-and-how-much-is-harness-17003604
Stable URL
/research/astra-s-computer-use-claims-how-much-is-model-and-how-much-is-harness-17003604
Version
v2
Coverage
2026-09-03
Records
29
Fields
7
Updated

Read the data

The records behind the figure

CSV JSON
Astra’s Computer-Use Claims: How Much Is Model, and How Much Is Harness? data records
EntityMetricValueUnitObservedSourceTransform
Agents' Last ExamAstra disclosed score59.3%percent2026-09-03https://openai.com/index/gpt-6-astra
Artificial Analysis Coding Agent Index v1.4Astra max in Codex67index points2026-09-03https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra
ScreenSpot-ProAstra no-tools score92.7%percent2026-09-03https://openai.com/index/gpt-6-astra
OSWorld 2.0 v2026.08.08 offline setAstra partial score72.6%percent2026-09-03https://openai.com/index/gpt-6-astra
DeepSWE v1.1Astra score74.1% at xhighpercent2026-09-03https://deepswe.datacurve.ai
FrontierCode 1.1 ExtendedAstra score64.5%percent2026-09-03https://openai.com/index/gpt-6-astra
FrontierCode 1.1 MainAstra score53.3%percent2026-09-03https://openai.com/index/gpt-6-astra
Internal Database Migration TasksAstra score63.9%percent2026-09-03https://openai.com/index/gpt-6-astra
Terminal-Bench 4.0Astra score57.9%percent2026-09-03https://openai.com/index/gpt-6-astra
Agents' Last ExamComplete requested protocol fieldsnonefields out of 92026-09-04https://openai.com/index/gpt-6-astraCounted only requested fields explicitly tied to the Astra result.
Artificial Analysis Coding Agent Index v1.4Complete requested protocol fieldsagent scaffold; retries; model settings; task split; evaluator versionfields out of 92026-09-04https://artificialanalysis.ai/methodology/coding-agents-benchmarkingCounted Codex identity, three attempts per task, max effort, the published 326-task component split and v1.4.
DeepSWE v1.1Complete requested protocol fieldsagent scaffold; tools; model settings; task split; evaluator versionfields out of 92026-09-04https://deepswe.datacurve.ai/blog/deepsweCounted mini-swe-agent, Bash, xhigh, the 113-task set and v1.1.
FrontierCode 1.1 ExtendedComplete requested protocol fieldstask split; evaluator versionfields out of 92026-09-04https://openai.com/index/gpt-6-astraCounted the Extended split label and version 1.1.
FrontierCode 1.1 MainComplete requested protocol fieldstask split; evaluator versionfields out of 92026-09-04https://openai.com/index/gpt-6-astraCounted the Main split label and version 1.1.
Internal Database Migration TasksComplete requested protocol fieldsnonefields out of 92026-09-04https://openai.com/index/gpt-6-astraCounted only requested fields explicitly tied to the result.
OSWorld 2.0 offlineComplete requested protocol fieldsbrowser policy; task split; evaluator versionfields out of 92026-09-04https://openai.com/index/gpt-6-astraCounted offline policy, offline split and v2026.08.08 evaluator identifier.
ScreenSpot-ProComplete requested protocol fieldstools: nonefields out of 92026-09-04https://openai.com/index/gpt-6-astraCounted the explicit no-tools label; general benchmark settings were not attributed to Astra without a run record.
Terminal-Bench 4.0Complete requested protocol fieldsevaluator versionfields out of 92026-09-04https://openai.com/index/gpt-6-astraCounted the 4.0 version identifier only.
Matched same-benchmark Astra raw-model versus harness comparisons0comparisons2026-09-04https://openai.com/index/gpt-6-astra
Results disclosing agent scaffold identity2 of 9percent of scoped results2026-09-04https://artificialanalysis.ai/methodology/coding-agents-benchmarking2 divided by 9 multiplied by 100.
Results disclosing a hard time or token budget0 of 9percent of scoped results2026-09-04https://openai.com/index/gpt-6-astraObserved runtime and token use were excluded because they are not caps.
Results disclosing all nine requested fields0 of 9percent of scoped results2026-09-04https://openai.com/index/gpt-6-astraCounted results with a completeness score of 9; none qualified.
Results disclosing complete browser/network policy1 of 9percent of scoped results2026-09-04https://openai.com/index/gpt-6-astra1 divided by 9 multiplied by 100; OSWorld's offline condition counted.
Results disclosing exact Astra model settings2 of 9percent of scoped results2026-09-04https://deepswe.datacurve.ai2 divided by 9 multiplied by 100; DeepSWE xhigh and Artificial Analysis max counted.
Results disclosing retries or attempts1 of 9percent of scoped results2026-09-04https://artificialanalysis.ai/methodology/coding-agents-benchmarking1 divided by 9 multiplied by 100.
Results disclosing state-management policy0 of 9percent of scoped results2026-09-04https://openai.com/index/gpt-6-astra0 divided by 9 multiplied by 100.
Results disclosing tool policy2 of 9percent of scoped results2026-09-04https://deepswe.datacurve.ai/blog/deepswe2 divided by 9 multiplied by 100; includes ScreenSpot-Pro's explicit no-tools condition.
Results identifying evaluator or benchmark version6 of 9percent of scoped results2026-09-04https://openai.com/index/gpt-6-astra6 divided by 9 multiplied by 100.
Results identifying task split5 of 9percent of scoped results2026-09-04https://openai.com/index/gpt-6-astra5 divided by 9 multiplied by 100.

Measurement technique

How to read this report

  1. 01Plan an evidence matrix with nine benchmark-result rows and nine protocol-field columns: scaffold, tools, state management, browser policy, retries, time/token budget, model settings, task split, and evaluator version.
  2. 02Count a field only when it is explicitly tied to the Astra result; do not treat generic model documentation, observed consumption, or general benchmark defaults as run-specific disclosure.
  3. 03Classify each result as a common-harness comparison, a model-plus-native-harness result, a no-tools result, or an insufficiently specified result.
  4. 04Keep benchmark scores separate from protocol completeness: a reported score does not establish transferability into another agent stack.
  5. 05Use benchmark and evaluator documentation to distinguish stated benchmark-wide procedures from procedures explicitly attributed to Astra’s run.

Sources

Evidence

7 publishers supporting 29 records. Expand a publisher to inspect its cited pages.

Next report / 01AI Model Economics Index All research reports