Open Models Are Catching the Frontier Faster. Benchmark Scores Aren’t the Whole Contest.

SemiAnalysis sees the lag from closed-model breakthroughs to open-model benchmark parity shrinking across three AI eras. But a comparable model score still leaves a harder product question: who can turn capability into dependable agent work?

By 4 min read
Open Models Are Catching the Frontier Faster. Benchmark Scores Aren’t the Whole Contest.
Open Models Are Catching the Frontier Faster. Benchmark Scores Aren’t the Whole Contest.

Listen to this story

The audio brief

About 1:49
0:001:49
Read transcript
Kimi K2.6 surpassed Opus 4.5 on SemiAnalysis’s benchmark suite just 4.8 months after the closed model appeared. GLM-5.2 passed GPT-5.2 in six months. That is the headline from SemiAnalysis’s cross-era comparison: open models are reaching closed-model benchmark parity faster each time the frontier resets. But the more important finding is that benchmark parity is arriving before product parity. The pattern spans three eras. In early scaling, the gap was wide: GPT-3.5 Turbo scored 75.7 on SemiAnalysis’s normalized composite, versus 39.9 for Llama 2 70B. In the reasoning era, the initial gap was 12.1 points, and R1-0528 closed it in 8.5 months. DeepSeek V3 later came close to GPT-4o, scoring 94.1 versus 95.5. Now the contest is about agents handling longer software, research, and knowledge-work tasks. SemiAnalysis used tests including Terminal-Bench 2.1, BrowseComp-Plus, τ³-banking, and DeepSWE. Meanwhile, OpenAI and Anthropic are releasing models about every 51 days, versus 120 days in the reasoning era. That can reset the target before an open model becomes a dependable product. The key distinction is the harness: the software and workflow layer that lets a model use tools and complete work. GPT-5.2 scored higher than Opus 4.5 on the suite, but SemiAnalysis judged Opus’s surrounding agent product more reliable. The constraint to watch is whether open releases can pair comparable scores with dependable harnesses before the next frontier reset.

Story brief

3 key points

SemiAnalysis’s historical comparison finds open models closing the gap with closed systems faster across three capability eras, but benchmark parity is arriving before product parity. Its latest examples include Kimi K2.6 surpassing Opus 4.5 in 4.8 months and GLM-5.2 surpassing GPT-5.2 in six. Meanwhile, OpenAI and Anthropic are releasing roughly every 51 days, potentially resetting the frontier faster than rivals...

  1. 01

    SemiAnalysis measured a 12.1-point initial reasoning-era gap versus 35.8 points at early scaling.

  2. 02

    DeepSeek V3 reached 94.1 versus GPT-4o’s 95.5 on SemiAnalysis’s normalized composite in December 2024.

  3. 03

    OpenAI and Anthropic averaged one model release every 51 days in the agentic era, versus 120 during reasoning.

In the reasoning and agentic eras, open AI models reached closed-model benchmark parity in months, according to SemiAnalysis’s cross-era analysis. The finding raises the pressure on frontier labs, but it does not settle the commercial contest: in the agentic era, the model is only one part of the product users experience.

The first chase began with a wide score gap

SemiAnalysis separates large-language-model progress into early scaling, reasoning and agentic eras. Its premise is that each frontier breakthrough resets the relevant tests, then competitors reproduce advances through methods including replication and distillation until open models narrow the gap.

In the early-scaling period, the firm’s normalized four-benchmark composite put GPT-3.5 Turbo at 75.7 and Llama 2 70B at 39.9. The score gap reflected an era defined by multiple-choice questions, word problems and narrowly scoped programming tasks, rather than the longer workflows that later became central.

The catch-up did not stop at the earlier frontier. SemiAnalysis says DeepSeek V3 matched GPT-4o in December 2024, with composite scores of 94.1 for DeepSeek V3 and 95.5 for GPT-4o. Those values are close, not identical, and the match reflects the firm’s composite framework.

Reasoning changed the measuring stick

The reasoning era began with OpenAI’s o1-preview in September 2024, which made older grade-school-style evaluations less useful for distinguishing leading systems. SemiAnalysis says the initial open-versus-closed composite gap was 12.1 points, far smaller than the 35.8-point gap it measured at the start of early scaling.

That smaller starting gap is the middle link in SemiAnalysis’s larger historical interpretation: open models have taken roughly half as long to catch the first closed model in each successive era. The pattern is not a forecast or a universal ranking. It depends on the firm’s model choices, dates and benchmark composites.

Agents move the contest beyond the model

The current agentic era shifts the test toward longer tasks in software engineering, research and knowledge work. SemiAnalysis used Terminal-Bench 2.1, BrowseComp-Plus, τ³-banking and DeepSWE, choosing newer benchmarks in part to limit memorization.

Frontier release cycles also accelerated. SemiAnalysis calculates that OpenAI and Anthropic released a model every 51 days on average during this era, compared with 213 days in early scaling and 120 days in reasoning. Faster releases can reset the reference point before an open rival has fully translated benchmark progress into a product.

SemiAnalysis argues that benchmark results do not automatically determine user experience. It found GPT-5.2 stronger than Opus 4.5 on its curated suite, while judging Opus 4.5’s surrounding agent product more reliable for users. The distinction is the harness: the software and workflow layer that lets a model use tools and complete work, rather than merely generate an answer.

A composite score has clear limits

SemiAnalysis normalized each era to its best result and averaged selected benchmarks with equal weights. It served open models using release-era vLLM versions, hardware and model-card sampling settings, while testing closed models through pinned API versions. That design aims for more consistent historical comparisons, not a universal capability ranking.

  • The composite is a curated, equal-weight benchmark measure rather than a direct measure of business value or day-to-day usability.
  • SemiAnalysis ran most scores with Prime Intellect’s evaluation stack, while drawing some values from Artificial Analysis and Datacurve’s DeepSWE leaderboard.
  • Public benchmarks can be optimized through training environments that closely mimic their tasks, limiting their usefulness as a proxy for real work.

The unresolved contest is whether open releases can pair comparable benchmark capability with dependable harnesses and product workflows while frontier labs introduce the next capability reset. SemiAnalysis’s history supports faster catch-up on its framework; it does not establish how long a product advantage can last.

Editorial analysis

Our Read

The strategic tension is not whether open models can reach selected benchmark thresholds; SemiAnalysis’s timeline says they increasingly can. The harder test is whether the same releases acquire the surrounding harnesses that shape reliable agent work. Watch whether benchmark leaders such as Kimi K2.6 and GLM-5.2 become available through products that can compete with established coding and workflow tools. A score above a closed rival would show capability catch-up on this framework. A usable model-plus-harness offering would test whether that catch-up changes the product market.

Citation desk / original work

Cite this

Permanent attributionView citation
Finding 01

The strategic tension is not whether open models can reach selected benchmark thresholds; SemiAnalysis’s timeline says they increasingly can.

/posts/open-models-are-catching-the-frontier-faster-benchmark-scores-aren-t-the-whole-contest#finding-1

Loading discussion...