Modelspublished

Qwen3.8-Max Leads Open-Weight Commerce Agent Bench by Two Passes, With Audit Limits

The benchmark checks whether agents leave commerce software and files in the required state, rather than simply narrate a workflow. But configuration choices and unavailable immutable result bundles limit what the narrow ranking can settle.

By 2 min read
Qwen3.8-Max Leads Open-Weight Commerce Agent Bench by Two Passes, With Audit Limits
Qwen3.8-Max Leads Open-Weight Commerce Agent Bench by Two Passes, With Audit Limits

Listen to this story

The audio brief

About 1:33
0:001:33
Read transcript
Qwen3.8-Max has taken a narrow lead in the published open-weight results for Commerce Agent Bench: 156 successful tasks, just two more than DeepSeek V4 Pro. But that lead is concentrated in one test harness, and the underlying records are not fully available for independent verification. Qwen passed 56 of 107 tasks in Accio, compared with 54 for DeepSeek. In OpenClaw, both models passed 47. In Pi, both passed 53. So the entire two-pass advantage came from Accio; it is not a broad lead across the benchmark. The comparison is also far behind Claude Opus 5, which recorded 191 aggregate passes. Even Claude’s best single-harness result was 66 out of 107, showing how difficult these tasks are. The benchmark covers 107 workflows across 14 offline commerce replicas, including product publishing, freight booking, payment operations, research, and spreadsheet creation. What makes the test meaningful is that an agent must leave software and files in the required final state. Describing a plausible workflow is not enough. At the same time, scores can move with routing, model snapshots, prompt adapters, retries, judge endpoints, and Qwen’s reasoning-content replay protocol. OpenClaw can be rerun, but Accio is reference-only. And the published task-level result bundles lack immutable public checksums. The key constraint, then, is whether this narrow ranking survives a fully reproducible run with the same configuration.

Story brief

3 key points

Commerce Agent Bench’s results show a narrow open-weight gap but a wide field-wide one: Qwen3.8-Max posted 156 successful tasks, two ahead of DeepSeek V4 Pro and 35 behind Claude Opus 5. Qwen and DeepSeek tied on OpenClaw and Pi, so Accio alone created the margin. The benchmark’s 107 state-checked tasks span 14 offline commerce replicas, but missing immutable task-level result archives and non-reproducible Accio...

  1. 01

    Qwen passed 56 Accio tasks, 47 OpenClaw tasks, and 53 Pi tasks; DeepSeek scored 54, 47, and 53.

  2. 02

    Claude Opus 5 reached 191 aggregate passes, but its best single-harness result was 66/107.

  3. 03

    Every task requires all checks to pass after execution; plausible action descriptions do not count.

Qwen3.8-Max holds a two-pass lead over DeepSeek V4 Pro in the published open-weight Commerce Agent Bench results, with 156 passes across three agent harnesses. The lead came from one harness, and the task-level records needed to fully reproduce the leaderboard are not publicly archived with immutable checksums.

Qwen3.8-Max completed 56 of 107 tasks through Accio, 47 through OpenClaw and 53 through Pi. DeepSeek V4 Pro recorded 54, 47 and 53 passes, respectively. The two-pass aggregate advantage therefore came entirely from Accio; the models tied in the other two harnesses.

It is not an overall lead. Claude Opus 5 led all evaluated models with 191 aggregate passes, 35 more than Qwen3.8-Max. Its best harness result was 66 of 107 tasks, leaving 41 tasks unfinished even at the published high-water mark.

Published open-weight results

01156 passes

Qwen3.8-Max

Qwen3.8-Max recorded 156 passes across the three published harness tables.

02154 passes

DeepSeek V4 Pro

DeepSeek V4 Pro recorded 154 passes in the same aggregate comparison.

Commerce Agent Bench contains 107 tasks spanning command-line, browser, file and document, and API or Model Context Protocol workflows. It covers product publishing, freight booking, storefront administration, payment operations, supplier analysis, web research and spreadsheet creation.

Each task starts in a fresh container within one of 14 offline software replicas. After an agent stops, the benchmark checks the resulting state of mock business services and any produced files. A task passes only if every required check passes, rather than because the agent described a plausible sequence of actions.

The materials identify several variables that can change a score: provider routing, model snapshots, prompt adapters, retry policies and judge endpoints. Qwen3.8-Max’s published rows also used a reasoning-content replay protocol, which carries its reasoning trace into later assistant messages and is described as necessary for comparable results.

The benchmark code is published under Apache 2.0 and its task data under CC BY 4.0, with commercial use permitted under those licenses. Yet Accio is reference-only and cannot be reproduced from a repository checkout, while OpenClaw is the shipped rerunnable harness. The published task-level result bundles are neither stored in Git nor available through immutable public archives with checksums, leaving full independent reproduction of the leaderboard unresolved.