Garp Independent AI & technology journalism
Sunday, September 27, 2026 Sign In · Join Subscribe
Latest Don’t be fooled by this summer of AI hype 

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  AI News  /  Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward

AI News

Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward

Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls…

OpenAI’s GPT-6 Astra is drawing contradictory benchmark verdicts. Epoch AI puts it out in front, while Artificial Analysis rates it no better than its predecessor.

The biggest surprise comes from ARC-AGI-3, where Astra works more efficiently than the average human for the first time. ARC Prize chief François Chollet calls the progress “2x faster” than he expected and is moving up his AGI forecast. Two independent labs each roll dozens of individual tests into a single overall score, but they reach opposite conclusions. Epoch AI combines more than 50 benchmarks and puts GPT-6 Astra clearly in first place with 169 points, ahead of 267 models. Artificial Analysis tests knowledge, coding, and text comprehension, and rates GPT-6 Astra at 61 points, exactly level with its predecessor and behind Claude Fable 5.1 at 66 points. Astra is clearly more expensive than its own predecessor. OpenAI charges two and a half times as much per unit of processed text, which makes a task cost roughly 75 percent more than it did with Sol. Compared with Anthropic, the picture flips. On coding tasks, Astra hits the same score as Claude Fable 5 according to Artificial Analysis, but costs less than half as much per task. The reason is how sparing the model is. It needs only a third of the compute steps Sol uses and a fifth of what Opus 5 uses. *GPT-6 Astra at xhigh reasoning effort; at max it hits 97.5 percent. ARC-AGI-1 is now considered largely saturated. On the Coding Agent Index, it reaches 67 points at roughly a third of Sol’s token usage, while Fable 5.1 leads with 70. The hallucination rate on AA-Omniscience drops from 92 to 51 percent.