Garp Independent AI & technology journalism
Friday, August 7, 2026 Sign In · Join Subscribe
Latest Defense tech Hadrian raises $1.37B at $8B valuation

AI news, research, models, robotics, chips, startups, and infrastructure coverage.

Updated daily

Home  /  AI News  /  Anthropic’s Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence

AI News

Anthropic’s Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence

Anthropic’s Opus 5 blows past Fable 5 and GPT-5.6 Sol on the…

The creators of the ARC-AGI benchmark say Anthropic’s Claude Opus 5 owes its massive lead on ARC-AGI-3 to genuinely better reasoning. The model scored 30.2 percent on ARC-AGI-3, making it the new leader.

The previous record was 7.8 percent, set by OpenAI’s GPT-5.6 Sol (Max). Opus 5 solved five previously unsolved environments, four of them at or above human level. That also puts it ahead of Anthropic’s “Fable-class” models, which hit around 20 percent according to ARC Prize. ARC Prize’s analysis credits the lead to stronger logical reasoning, “which enables more autonomous exploration, planning, and execution across unfamiliar environments.” During testing, Opus 5 also showed behavior that researchers hadn’t seen from a model before. It translated tasks into algebraic notation and independently formulated reflection equations for the first time.Ad Six of the 25 public demo environments have now been solved. The full results, replays, and benchmarking code are publicly available. On the older ARC-AGI-2 benchmark, Opus 5 scores 90.4 percent, and it reaches 97.5 percent on ARC-AGI-1. Both results match previous top scores, though at slightly higher costs, according to ARC Prize.AdDEC_D_Incontent-1 ARC-AGI-3 measures how well AI models solve new tasks they didn’t encounter during training, including ones humans can usually handle with ease. The current version works like a game. The model must infer the rules of an interactive environment, plan its actions, and carry them out step by step. This tests general reasoning rather than stored knowledge. Some AI systems may have already passed the benchmark, but they rely on extra software known as a harness. Official scores count only the language model’s own performance. ARC Prize argues that future AGI systems shouldn’t need outside help to solve new tasks. Opus 5 would likely score even higher if used within Claude Code.Ad Anthropic hasn’t explained the gain, but targeted data labeling and reinforcement learning are plausible factors.