{"title":"Astra’s Computer-Use Claims: How Much Is Model, and How Much Is Harness?","description":"Astra may be highly capable, but the disclosed results do not consistently isolate model performance from the surrounding agent system.","dataset_id":"spd:astra-s-computer-use-claims-how-much-is-model-and-how-much-is-harness-17003604","canonical_url":"https://superpowerdaily.com/research/astra-s-computer-use-claims-how-much-is-model-and-how-much-is-harness-17003604","version_url":"https://superpowerdaily.com/research/astra-s-computer-use-claims-how-much-is-model-and-how-much-is-harness-17003604/versions/v2","version":"v2","snapshot_hash":"fa14e8ddce696f35b2508317841e09b2cac04138bf3e1e0b6fd73459f4f65c57","date_created":"2026-09-05T00:31:03.248Z","date_modified":"2026-09-06T20:33:34.117Z","license":{"name":"Superpower Daily data reuse terms","url":"https://superpowerdaily.com/terms"},"license_url":"https://superpowerdaily.com/terms","temporal_coverage":"2026-09-03","coverage_note":"The ledger covers nine launch-table or independently reported Astra rows and excludes ARC-AGI-3. Field coverage is highest for evaluator or benchmark version (6 of 9) and task split (5 of 9), while state management and hard time/token budgets are disclosed for none of the nine rows.","measurement_technique":["Plan an evidence matrix with nine benchmark-result rows and nine protocol-field columns: scaffold, tools, state management, browser policy, retries, time/token budget, model settings, task split, and evaluator version.","Count a field only when it is explicitly tied to the Astra result; do not treat generic model documentation, observed consumption, or general benchmark defaults as run-specific disclosure.","Classify each result as a common-harness comparison, a model-plus-native-harness result, a no-tools result, or an insufficiently specified result.","Keep benchmark scores separate from protocol completeness: a reported score does not establish transferability into another agent stack.","Use benchmark and evaluator documentation to distinguish stated benchmark-wide procedures from procedures explicitly attributed to Astra’s run."],"methodology":["Plan an evidence matrix with nine benchmark-result rows and nine protocol-field columns: scaffold, tools, state management, browser policy, retries, time/token budget, model settings, task split, and evaluator version.","Count a field only when it is explicitly tied to the Astra result; do not treat generic model documentation, observed consumption, or general benchmark defaults as run-specific disclosure.","Classify each result as a common-harness comparison, a model-plus-native-harness result, a no-tools result, or an insufficiently specified result.","Keep benchmark scores separate from protocol completeness: a reported score does not establish transferability into another agent stack.","Use benchmark and evaluator documentation to distinguish stated benchmark-wide procedures from procedures explicitly attributed to Astra’s run."],"metrics":[{"label":"Verified observations","value":"29","detail":"17 measured fields"},{"label":"Supported claims","value":"9","detail":"9 material findings"},{"label":"Cited sources","value":"11","detail":"11 primary or authoritative"},{"label":"Research score","value":"83","detail":"Automated topic and evidence score"}],"columns":[{"key":"entity","label":"Entity"},{"key":"metric","label":"Metric"},{"key":"value","label":"Value"},{"key":"unit","label":"Unit"},{"key":"observed","label":"Observed"},{"key":"source","label":"Source"},{"key":"transform","label":"Transform"}],"data":[{"unit":"percent","value":"59.3%","entity":"Agents' Last Exam","metric":"Astra disclosed score","source":"https://openai.com/index/gpt-6-astra","observed":"2026-09-03","transform":null},{"unit":"index points","value":"67","entity":"Artificial Analysis Coding Agent Index v1.4","metric":"Astra max in Codex","source":"https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra","observed":"2026-09-03","transform":null},{"unit":"percent","value":"92.7%","entity":"ScreenSpot-Pro","metric":"Astra no-tools score","source":"https://openai.com/index/gpt-6-astra","observed":"2026-09-03","transform":null},{"unit":"percent","value":"72.6%","entity":"OSWorld 2.0 v2026.08.08 offline set","metric":"Astra partial score","source":"https://openai.com/index/gpt-6-astra","observed":"2026-09-03","transform":null},{"unit":"percent","value":"74.1% at xhigh","entity":"DeepSWE v1.1","metric":"Astra score","source":"https://deepswe.datacurve.ai","observed":"2026-09-03","transform":null},{"unit":"percent","value":"64.5%","entity":"FrontierCode 1.1 Extended","metric":"Astra score","source":"https://openai.com/index/gpt-6-astra","observed":"2026-09-03","transform":null},{"unit":"percent","value":"53.3%","entity":"FrontierCode 1.1 Main","metric":"Astra score","source":"https://openai.com/index/gpt-6-astra","observed":"2026-09-03","transform":null},{"unit":"percent","value":"63.9%","entity":"Internal Database Migration Tasks","metric":"Astra score","source":"https://openai.com/index/gpt-6-astra","observed":"2026-09-03","transform":null},{"unit":"percent","value":"57.9%","entity":"Terminal-Bench 4.0","metric":"Astra score","source":"https://openai.com/index/gpt-6-astra","observed":"2026-09-03","transform":null},{"unit":"fields out of 9","value":"none","entity":"Agents' Last Exam","metric":"Complete requested protocol fields","source":"https://openai.com/index/gpt-6-astra","observed":"2026-09-04","transform":"Counted only requested fields explicitly tied to the Astra result."},{"unit":"fields out of 9","value":"agent scaffold; retries; model settings; task split; evaluator version","entity":"Artificial Analysis Coding Agent Index v1.4","metric":"Complete requested protocol fields","source":"https://artificialanalysis.ai/methodology/coding-agents-benchmarking","observed":"2026-09-04","transform":"Counted Codex identity, three attempts per task, max effort, the published 326-task component split and v1.4."},{"unit":"fields out of 9","value":"agent scaffold; tools; model settings; task split; evaluator version","entity":"DeepSWE v1.1","metric":"Complete requested protocol fields","source":"https://deepswe.datacurve.ai/blog/deepswe","observed":"2026-09-04","transform":"Counted mini-swe-agent, Bash, xhigh, the 113-task set and v1.1."},{"unit":"fields out of 9","value":"task split; evaluator version","entity":"FrontierCode 1.1 Extended","metric":"Complete requested protocol fields","source":"https://openai.com/index/gpt-6-astra","observed":"2026-09-04","transform":"Counted the Extended split label and version 1.1."},{"unit":"fields out of 9","value":"task split; evaluator version","entity":"FrontierCode 1.1 Main","metric":"Complete requested protocol fields","source":"https://openai.com/index/gpt-6-astra","observed":"2026-09-04","transform":"Counted the Main split label and version 1.1."},{"unit":"fields out of 9","value":"none","entity":"Internal Database Migration Tasks","metric":"Complete requested protocol fields","source":"https://openai.com/index/gpt-6-astra","observed":"2026-09-04","transform":"Counted only requested fields explicitly tied to the result."},{"unit":"fields out of 9","value":"browser policy; task split; evaluator version","entity":"OSWorld 2.0 offline","metric":"Complete requested protocol fields","source":"https://openai.com/index/gpt-6-astra","observed":"2026-09-04","transform":"Counted offline policy, offline split and v2026.08.08 evaluator identifier."},{"unit":"fields out of 9","value":"tools: none","entity":"ScreenSpot-Pro","metric":"Complete requested protocol fields","source":"https://openai.com/index/gpt-6-astra","observed":"2026-09-04","transform":"Counted the explicit no-tools label; general benchmark settings were not attributed to Astra without a run record."},{"unit":"fields out of 9","value":"evaluator version","entity":"Terminal-Bench 4.0","metric":"Complete requested protocol fields","source":"https://openai.com/index/gpt-6-astra","observed":"2026-09-04","transform":"Counted the 4.0 version identifier only."},{"unit":"comparisons","value":"0","entity":null,"metric":"Matched same-benchmark Astra raw-model versus harness comparisons","source":"https://openai.com/index/gpt-6-astra","observed":"2026-09-04","transform":null},{"unit":"percent of scoped results","value":"2 of 9","entity":null,"metric":"Results disclosing agent scaffold identity","source":"https://artificialanalysis.ai/methodology/coding-agents-benchmarking","observed":"2026-09-04","transform":"2 divided by 9 multiplied by 100."},{"unit":"percent of scoped results","value":"0 of 9","entity":null,"metric":"Results disclosing a hard time or token budget","source":"https://openai.com/index/gpt-6-astra","observed":"2026-09-04","transform":"Observed runtime and token use were excluded because they are not caps."},{"unit":"percent of scoped results","value":"0 of 9","entity":null,"metric":"Results disclosing all nine requested fields","source":"https://openai.com/index/gpt-6-astra","observed":"2026-09-04","transform":"Counted results with a completeness score of 9; none qualified."},{"unit":"percent of scoped results","value":"1 of 9","entity":null,"metric":"Results disclosing complete browser/network policy","source":"https://openai.com/index/gpt-6-astra","observed":"2026-09-04","transform":"1 divided by 9 multiplied by 100; OSWorld's offline condition counted."},{"unit":"percent of scoped results","value":"2 of 9","entity":null,"metric":"Results disclosing exact Astra model settings","source":"https://deepswe.datacurve.ai","observed":"2026-09-04","transform":"2 divided by 9 multiplied by 100; DeepSWE xhigh and Artificial Analysis max counted."},{"unit":"percent of scoped results","value":"1 of 9","entity":null,"metric":"Results disclosing retries or attempts","source":"https://artificialanalysis.ai/methodology/coding-agents-benchmarking","observed":"2026-09-04","transform":"1 divided by 9 multiplied by 100."},{"unit":"percent of scoped results","value":"0 of 9","entity":null,"metric":"Results disclosing state-management policy","source":"https://openai.com/index/gpt-6-astra","observed":"2026-09-04","transform":"0 divided by 9 multiplied by 100."},{"unit":"percent of scoped results","value":"2 of 9","entity":null,"metric":"Results disclosing tool policy","source":"https://deepswe.datacurve.ai/blog/deepswe","observed":"2026-09-04","transform":"2 divided by 9 multiplied by 100; includes ScreenSpot-Pro's explicit no-tools condition."},{"unit":"percent of scoped results","value":"6 of 9","entity":null,"metric":"Results identifying evaluator or benchmark version","source":"https://openai.com/index/gpt-6-astra","observed":"2026-09-04","transform":"6 divided by 9 multiplied by 100."},{"unit":"percent of scoped results","value":"5 of 9","entity":null,"metric":"Results identifying task split","source":"https://openai.com/index/gpt-6-astra","observed":"2026-09-04","transform":"5 divided by 9 multiplied by 100."}],"sources":[{"url":"https://artificialanalysis.ai/agents/coding-agents/comparisons/codex-vs-muse-code","name":"Artificial Analysis","title":"Codex vs Muse Code: Coding Agent Comparison","records":0},{"url":"https://artificialanalysis.ai/methodology/coding-agents-benchmarking","name":"Artificial Analysis","title":"Coding Agent Index Methodology | Artificial Analysis","records":3},{"url":"https://deepswe.datacurve.ai/blog/deepswe","name":"Datacurve","title":"deepswe.datacurve.ai","records":2},{"url":"https://deepswe.datacurve.ai","name":"Datacurve","title":"deepswe.datacurve.ai","records":2},{"url":"https://deepswe.datacurve.ai/changelog","name":"Datacurve","title":"deepswe.datacurve.ai","records":0},{"url":"https://github.com/xlang-ai/OSWorld","name":"XLANG Lab","title":"GitHub - xlang-ai/OSWorld: [NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments","records":0},{"url":"https://openai.com/index/gpt-6-astra","name":"OpenAI","title":"GPT-6 Astra: A new generation of intelligence","records":21},{"url":"https://developers.openai.com/api/docs/models/gpt-6-astra","name":"OpenAI","title":"GPT-6 Astra Model | OpenAI API","records":0},{"url":"https://gui-agent.github.io/grounding-leaderboard","name":"ScreenSpot-Pro authors","title":"gui-agent.github.io","records":0},{"url":"https://osworld-v2.xlang.ai","name":"XLANG Lab","title":"osworld-v2.xlang.ai","records":0},{"url":"https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra","name":"Artificial Analysis","title":"Benchmarking GPT-6 Astra","records":1}],"provenance":{"publisher":"Superpower Daily","source_count":11,"source_urls":["https://artificialanalysis.ai/agents/coding-agents/comparisons/codex-vs-muse-code","https://artificialanalysis.ai/methodology/coding-agents-benchmarking","https://deepswe.datacurve.ai/blog/deepswe","https://deepswe.datacurve.ai","https://deepswe.datacurve.ai/changelog","https://github.com/xlang-ai/OSWorld","https://openai.com/index/gpt-6-astra","https://developers.openai.com/api/docs/models/gpt-6-astra","https://gui-agent.github.io/grounding-leaderboard","https://osworld-v2.xlang.ai","https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra"],"methodology":["Plan an evidence matrix with nine benchmark-result rows and nine protocol-field columns: scaffold, tools, state management, browser policy, retries, time/token budget, model settings, task split, and evaluator version.","Count a field only when it is explicitly tied to the Astra result; do not treat generic model documentation, observed consumption, or general benchmark defaults as run-specific disclosure.","Classify each result as a common-harness comparison, a model-plus-native-harness result, a no-tools result, or an insufficiently specified result.","Keep benchmark scores separate from protocol completeness: a reported score does not establish transferability into another agent stack.","Use benchmark and evaluator documentation to distinguish stated benchmark-wide procedures from procedures explicitly attributed to Astra’s run."],"snapshot_hash":"fa14e8ddce696f35b2508317841e09b2cac04138bf3e1e0b6fd73459f4f65c57"},"distributions":{"csv":"https://superpowerdaily.com/api/research/astra-s-computer-use-claims-how-much-is-model-and-how-much-is-harness-17003604/versions/v2?format=csv","json":"https://superpowerdaily.com/api/research/astra-s-computer-use-claims-how-much-is-model-and-how-much-is-harness-17003604/versions/v2?format=json"}}