The ledger covers nine launch-table or independently reported Astra rows and excludes ARC-AGI-3. Field coverage is highest for evaluator or benchmark version (6 of 9) and task split (5 of 9), while state management and hard time/token budgets are disclosed for none of the nine rows.
2 divided by 9 multiplied by 100; includes ScreenSpot-Pro's explicit no-tools condition.
—
Results identifying evaluator or benchmark version
6 of 9
percent of scoped results
2026-09-04
https://openai.com/index/gpt-6-astra
6 divided by 9 multiplied by 100.
—
Results identifying task split
5 of 9
percent of scoped results
2026-09-04
https://openai.com/index/gpt-6-astra
5 divided by 9 multiplied by 100.
Measurement technique
How to read this report
01Plan an evidence matrix with nine benchmark-result rows and nine protocol-field columns: scaffold, tools, state management, browser policy, retries, time/token budget, model settings, task split, and evaluator version.
02Count a field only when it is explicitly tied to the Astra result; do not treat generic model documentation, observed consumption, or general benchmark defaults as run-specific disclosure.
03Classify each result as a common-harness comparison, a model-plus-native-harness result, a no-tools result, or an insufficiently specified result.
04Keep benchmark scores separate from protocol completeness: a reported score does not establish transferability into another agent stack.
05Use benchmark and evaluator documentation to distinguish stated benchmark-wide procedures from procedures explicitly attributed to Astra’s run.
Sources
Evidence
7 publishers supporting 29 records. Expand a publisher to inspect its cited pages.