Artificial Intelligence
Benchmarks split on GPT-6 Astra
Independent benchmark suites gave GPT-6 Astra different scores, while ARC Prize said its ARC-AGI-3 result beat human move efficiency and led François Chollet to bring forward his forecast.
- Epoch AI put GPT-6 Astra first on 169 points across more than 50 benchmarks, while Artificial Analysis gave it 61 points, level with GPT-5.6 Sol and below Claude Fable 5.1.
- ARC Prize said GPT-6 Astra scored 62.7 per cent on ARC-AGI-3 at about $26,000, versus 7.78 per cent for GPT-5.6 Sol and 30.16 per cent for Claude Opus 5.
- OpenAI's 99.9 per cent figure for ARC-AGI-3 came from a different harness; ARC Prize said the comparable setup ran about 3.66 times faster and used 49 per cent fewer tokens on the pairs both setups solved.
- On FrontierMath Erdős, Epoch AI said Astra was the only model to solve two of 68 open Erdős problems with Lean-verified proofs, while three extra solutions from non-standard runs did not count.
- François Chollet said the result was not proof of AGI, but that progress had arrived about twice as fast as he expected and that ARC-AGI-4 was due in the first quarter of 2027.
Epoch AI and Artificial Analysis reached different aggregate scores for GPT-6 Astra. Epoch AI combined more than 50 benchmarks and placed it first on 169 points, while Artificial Analysis scored it at 61 points, equal to GPT-5.6 Sol and below Claude Fable 5.1.
The article said OpenAI charged two and a half times as much per unit of processed text for Astra as for Sol, making one task about 75 per cent more expensive. It also said Artificial Analysis found Astra tied Claude Fable 5 on coding while costing less than half as much per task because it used fewer compute steps.
On FrontierMath Erdős, Epoch AI said Astra was the only model to solve two of 68 open Erdős problems with Lean-verified proofs at $300 per attempt. It said three further solutions came from extra runs that used more than $220,000 of compute and were left out of the score.
ARC Prize said Astra reached 62.7 per cent on ARC-AGI-3 with its internal harness, compared with 7.78 per cent for GPT-5.6 Sol and 30.16 per cent for Claude Opus 5. OpenAI's 99.9 per cent figure came from a different harness, and ARC Prize said that setup ran faster and used fewer tokens on the cases both setups solved.
ARC Prize said Astra used fewer moves than the median human solver on 96 per cent of levels and sometimes built its own shorthand notes or small code tools. Chollet said the result did not prove AGI, but that progress had come about twice as fast as he expected and that ARC-AGI-4 was due in the first quarter of 2027.
Named in this story
People
- François Chollet
- ARC Prize chief who said the result was not proof of AGI
Companies
- OpenAI
- made GPT-6 Astra and set its pricing
Organisations
- Epoch AI
- combined more than 50 benchmarks and ranked Astra first
- Artificial Analysis
- scored Astra alongside Sol and behind Claude Fable 5.1
- ARC Prize
- ran ARC-AGI-3 and compared harness results with OpenAI
Products and systems
- GPT-6 Astra
- the model measured across the benchmarks
- Claude Fable 5.1
- ranked ahead of Astra in Artificial Analysis's score
- Claude Fable 5
- matched Astra on coding in Artificial Analysis
- Claude Opus 5
- appeared in the ARC-AGI-3 comparison and coding discussion
- GPT-5.6 Sol
- Astra's predecessor in the comparisons
- ARC-AGI-3
- the benchmark where Astra's human-efficiency result was reported
- ARC-AGI-4
- the next benchmark planned for release in early 2027
- FrontierMath Erdős
- the math benchmark where Astra solved two open problems
- PRO-LONG
- the third-party agent framework used as a red-teaming partner
How the source tells it
The source reads like benchmark reporting, but it leans on surprise, speed and first-time comparison language before ending with a caution that the result is not proof of AGI.
- novelty hype the headline and repeated first-time framing around the benchmark results
- progress acceleration comparative score language and the forecast being moved forward because progress was said to be faster
- qualification an explicit disclaimer that the result was not proof of AGI and that the benchmark was limited