Artificial Intelligence
Open models caught up faster in each era
SemiAnalysis compared open and closed frontier models across three LLM eras and said the time for open systems to close the first closed-model lead roughly halved each time.
- SemiAnalysis divided LLM development into early scaling, reasoning and agentic eras and scored selected models against curated benchmark sets.
- It said open-source models took about half as long in each era to catch the first closed model.
- It said DeepSeek V3 matched GPT-4o in December 2024, and R1-0528 closed the reasoning-era gap in May 2025.
- It said Kimi K2.6 and GLM-5.2 later overtook Opus 4.5 and GPT-5.2 within months.
SemiAnalysis said it compared open and closed models across three eras of LLM development: early scaling, reasoning and agentic work. It said the scores came from curated benchmark sets, with most results produced by its own evaluation setup.
In the early scaling era, it said Llama-2-70B was the first open model to move near the frontier, although it still trailed the leading closed model. It said later open releases closed that gap to the then-leading closed systems.
For the reasoning era, it said o1-preview reset which benchmarks mattered and that DeepSeek R1 reduced the opening gap. It said R1-0528 closed that gap in May 2025.
In the agentic era, it said terminal and browser tasks had become the relevant tests for long-horizon work. It said Kimi K2.6 and GLM-5.2 overtook Opus 4.5 and GPT-5.2 within months, and it warned that benchmarks did not capture every part of real work.
Named in this story
People
- Noam Brown
- was cited as a hire that pointed OpenAI towards reasoning
- Florian Brand
- helped select benchmarks and check the evals
Companies
- Anthropic
- was one of the frontier labs compared in the analysis
- OpenAI
- was one of the closed-model labs compared in the analysis
- DeepSeek
- released the open models R1, V3 and R1-0528
Organisations
- SemiAnalysis
- published the analysis and the benchmark comparison
Products and systems
- Llama-2-70B
- was the first open model to approach the frontier in early scaling
- GPT-3.5 Turbo
- was the closed baseline in the early-scaling comparison
- GPT-4o
- was the last frontier model matched in the early-scaling era
- DeepSeek V3
- matched GPT-4o in December 2024
- o1-preview
- opened the reasoning era
- DeepSeek R1
- narrowed the initial reasoning-era gap
- R1-0528
- closed the reasoning-era gap in May 2025
- Opus 4.5
- served as the agentic-era frontier comparator
- GPT-5.2
- was OpenAI's agentic-era flagship in the comparison
- Kimi K2.6
- surpassed Opus 4.5 in the agentic era
- GLM-5.2
- cleared GPT-5.2 in the agentic era
How the source tells it
The register was analytical but promotional, with era-setting language, market talk and repeated self-marketing alongside the benchmark claims.
- hype and novelty era language, breakthrough framing and celebratory adjectives used to make each model wave feel historic
- threat inflation competition was recast as a margin threat through FUD language and disastrous worst-case claims
- vendor boosterism the piece promoted its own numbers and products while contrasting them favourably with others