Open Models Are Catching the Frontier Faster. Benchmark Scores Aren’t the Whole Contest.
SemiAnalysis sees the lag from closed-model breakthroughs to open-model benchmark parity shrinking across three AI eras. But a comparable model score still leaves a harder product question: who can turn capability into dependable agent work?
Listen to this story
The audio brief
Story brief
3 key pointsSemiAnalysis’s historical comparison finds open models closing the gap with closed systems faster across three capability eras, but benchmark parity is arriving before product parity. Its latest examples include Kimi K2.6 surpassing Opus 4.5 in 4.8 months and GLM-5.2 surpassing GPT-5.2 in six. Meanwhile, OpenAI and Anthropic are releasing roughly every 51 days, potentially resetting the frontier faster than rivals...
- 01
SemiAnalysis measured a 12.1-point initial reasoning-era gap versus 35.8 points at early scaling.
- 02
DeepSeek V3 reached 94.1 versus GPT-4o’s 95.5 on SemiAnalysis’s normalized composite in December 2024.
- 03
OpenAI and Anthropic averaged one model release every 51 days in the agentic era, versus 120 during reasoning.
In the reasoning and agentic eras, open AI models reached closed-model benchmark parity in months, according to SemiAnalysis’s cross-era analysis. The finding raises the pressure on frontier labs, but it does not settle the commercial contest: in the agentic era, the model is only one part of the product users experience.
The first chase began with a wide score gap
SemiAnalysis separates large-language-model progress into early scaling, reasoning and agentic eras. Its premise is that each frontier breakthrough resets the relevant tests, then competitors reproduce advances through methods including replication and distillation until open models narrow the gap.
In the early-scaling period, the firm’s normalized four-benchmark composite put GPT-3.5 Turbo at 75.7 and Llama 2 70B at 39.9. The score gap reflected an era defined by multiple-choice questions, word problems and narrowly scoped programming tasks, rather than the longer workflows that later became central.
The catch-up did not stop at the earlier frontier. SemiAnalysis says DeepSeek V3 matched GPT-4o in December 2024, with composite scores of 94.1 for DeepSeek V3 and 95.5 for GPT-4o. Those values are close, not identical, and the match reflects the firm’s composite framework.
Reasoning changed the measuring stick
The reasoning era began with OpenAI’s o1-preview in September 2024, which made older grade-school-style evaluations less useful for distinguishing leading systems. SemiAnalysis says the initial open-versus-closed composite gap was 12.1 points, far smaller than the 35.8-point gap it measured at the start of early scaling.
That smaller starting gap is the middle link in SemiAnalysis’s larger historical interpretation: open models have taken roughly half as long to catch the first closed model in each successive era. The pattern is not a forecast or a universal ranking. It depends on the firm’s model choices, dates and benchmark composites.
Agents move the contest beyond the model
The current agentic era shifts the test toward longer tasks in software engineering, research and knowledge work. SemiAnalysis used Terminal-Bench 2.1, BrowseComp-Plus, τ³-banking and DeepSWE, choosing newer benchmarks in part to limit memorization.
Frontier release cycles also accelerated. SemiAnalysis calculates that OpenAI and Anthropic released a model every 51 days on average during this era, compared with 213 days in early scaling and 120 days in reasoning. Faster releases can reset the reference point before an open rival has fully translated benchmark progress into a product.
SemiAnalysis argues that benchmark results do not automatically determine user experience. It found GPT-5.2 stronger than Opus 4.5 on its curated suite, while judging Opus 4.5’s surrounding agent product more reliable for users. The distinction is the harness: the software and workflow layer that lets a model use tools and complete work, rather than merely generate an answer.
A composite score has clear limits
SemiAnalysis normalized each era to its best result and averaged selected benchmarks with equal weights. It served open models using release-era vLLM versions, hardware and model-card sampling settings, while testing closed models through pinned API versions. That design aims for more consistent historical comparisons, not a universal capability ranking.
- The composite is a curated, equal-weight benchmark measure rather than a direct measure of business value or day-to-day usability.
- SemiAnalysis ran most scores with Prime Intellect’s evaluation stack, while drawing some values from Artificial Analysis and Datacurve’s DeepSWE leaderboard.
- Public benchmarks can be optimized through training environments that closely mimic their tasks, limiting their usefulness as a proxy for real work.
The unresolved contest is whether open releases can pair comparable benchmark capability with dependable harnesses and product workflows while frontier labs introduce the next capability reset. SemiAnalysis’s history supports faster catch-up on its framework; it does not establish how long a product advantage can last.
Editorial analysis
Our Read
The strategic tension is not whether open models can reach selected benchmark thresholds; SemiAnalysis’s timeline says they increasingly can. The harder test is whether the same releases acquire the surrounding harnesses that shape reliable agent work. Watch whether benchmark leaders such as Kimi K2.6 and GLM-5.2 become available through products that can compete with established coding and workflow tools. A score above a closed rival would show capability catch-up on this framework. A usable model-plus-harness offering would test whether that catch-up changes the product market.
Citation desk / original work
Cite this
Citation desk / original work
Cite this
The strategic tension is not whether open models can reach selected benchmark thresholds; SemiAnalysis’s timeline says they increasingly can.
/posts/open-models-are-catching-the-frontier-faster-benchmark-scores-aren-t-the-whole-contest#finding-1
Loading discussion...
Reader comments
Newest comments first. Replies stay oldest first.