Mozilla Finds Chinese Open Models Are Within 4.4 Months of U.S. Frontier AI
Mozilla’s latest analysis argues that the expensive frontier-model premium is increasingly limited to difficult, long-running work—while the market for open models raises a new concentration problem.
Listen to this story
The audio brief
Story brief
3 key pointsMozilla’s latest open-weight AI analysis argues that Chinese models now deliver near-frontier capability at a fraction of U.S. closed-model pricing, but the advantage has limits. Its comparisons put leading open systems roughly 4.4 months behind, with the best closed model reliably handling 12-hour tasks versus seven hours for the best open model. Mozilla recommends open models for most routine organizational work,...
- 01
Kimi K3 scored three points below Anthropic’s Fable 5 on Artificial Analysis’s index while costing about 30% as much.
- 02
Z.ai’s GLM 5.2 came within one point of Claude Opus 4.7 and 4.8 on Terminal-Bench 2.1 at roughly one-fifth the per-task cost.
- 03
METR found the best closed model handled tasks 1.7 times longer: about 12 hours versus seven for the best open model.
The promise of closed frontier AI has been clear: pay more for the best capability. Mozilla’s new analysis suggests that advantage has become narrower. It puts leading Chinese open-weight models about 4.4 months behind leading U.S. closed models, a gap that can matter for demanding work but not necessarily for the bulk of everyday AI tasks.
That is Mozilla’s central recommendation: use open models as the default for most work, then pay for closed systems when a particular job calls for their strengths. CTO Raffi Krikorian identified expert professional work, high-intensity retrieval, and long-context tasks as the areas where closed models still earn their premium.
A cost advantage with important boundaries
Mozilla’s comparison pairs Moonshot AI’s Kimi K3 with Anthropic’s Fable 5 on the Artificial Analysis Intelligence Index. The report says Kimi K3 scored three points lower while costing about 30% as much. That does not mean the models are interchangeable in every setting; it is one composite benchmark and price comparison, not a complete measure of production reliability.
Vals AI ran models through its own neutral software harness—the layer that gives a model tools and memory for agent-like work—on Terminal-Bench 2.1. Z.ai’s open-weight GLM 5.2 finished within one point of Anthropic’s Claude Opus 4.7 and 4.8, at roughly one-fifth of the per-task cost.
The shrinking window for a premium model
Mozilla also draws the line through task duration. A METR time-horizon comparison found the best closed model could reliably handle tasks 1.7 times longer than the best open model. Krikorian described the current divide as roughly seven-hour jobs for the leading open model and 12-hour jobs for the closed frontier.
- Tasks under eight hours could generally be handled by either type of model, according to Mozilla.
- The narrower eight-to-12-hour band is where Mozilla says closed models currently have an edge.
- Mozilla says no models generally complete tasks longer than 12 hours reliably.
The distinction matters because a benchmark score alone cannot capture a model’s surrounding product. Closed-model providers often supply their own harnesses, which can improve performance. Many organizations also lack the staff to operate open-weight models well, while closed offerings arrive with support, compliance packaging, and clearer accountability, Krikorian said.
Openness has its own concentration risk
Mozilla warns that the strongest open models are concentrated in China, while leading closed frontier systems come from U.S. companies. Eight of the 10 models ranked by token volume on OpenRouter in August provided open weights, according to the report, underscoring their reach without resolving who will control the ecosystem.
That reach has not translated into comparable sales: a Linux Foundation paper cited by Mozilla found open models received 4% of overall model revenue, compared with 96% for closed models, using data from May through September 2025. Mozilla’s earlier inaugural report, based on new analysis and a survey of more than 950 developers, likewise found that open models were widely used but less often deployed in production than closed ones.
The practical outcome is a more selective buying decision. The evidence supports a cheaper open default for many routine workloads, not a declaration that closed models no longer matter. It also leaves a strategic question: if organizations move toward open weights for cost and control, whether they will accept an ecosystem whose current leaders are concentrated in one country.
Sources
- blog.mozilla.orgMozilla’s Inaugural ‘State of Open Source AI’ Report Is Here | The Mozilla Blog
- arstechnica.comExclusive: Paying for frontier AI models buys 4-month head start at 5x the cost
Loading discussion...
Reader comments
Newest comments first. Replies stay oldest first.