Anthropic's Claude Opus 5 has achieved a significant milestone on the ARC-AGI-3 benchmark, which is designed to measure real intelligence. The model scored 30.2 percent, surpassing the previous record of 7.8 percent set by OpenAI's GPT-5.6 Sol (Max). This marks a substantial leap in performance, according to the ARC Prize team, which attributes the lead to stronger logical reasoning that enables more autonomous exploration and planning in unfamiliar environments. Opus 5 also solved five previously unsolved environments, with four of them at or above human level. This places it ahead of Anthropic's Fable-class models, which scored around 20 percent on the same benchmark. The model's ability to translate tasks into algebraic notation and independently formulate reflection equations represents new behaviors not previously observed in AI models. Source: thedecoder

The ARC-AGI-3 benchmark evaluates how well AI models can solve new tasks they haven't encountered during training, including those humans can handle with ease. The benchmark operates like a game, requiring models to infer rules, plan actions, and execute them step by step. This tests general reasoning rather than stored knowledge. While some AI systems may have passed the benchmark, they often rely on external software known as a harness. Official scores only reflect the language model's own performance, according to ARC Prize. Opus 5's performance on the older ARC-AGI-2 benchmark was 90.4 percent, and it reached 97.5 percent on ARC-AGI-1, matching previous top scores but at slightly higher costs. Source: thedecoder

Independent tests suggest narrower gains for Opus 5, though Anthropic has not explained the improvement. Targeted data labeling and reinforcement learning are plausible factors. Unlike earlier models, Opus 5 was developed after the ARC-AGI-3 benchmark became public, allowing Anthropic to focus on its skills and puzzle formats. However, the company did not train on the exact tasks. Annotators may have labeled reasoning traces, useful actions, and recovery steps from similar puzzles. Reinforcement learning could have rewarded exploration, planning, rule discovery, and self-correction. Tests on Guanghan Ning's private Witness benchmark showed Opus 5 scored 43.4, statistically tying Kimi K3 and Fable 5, but improving far less than on ARC-AGI-3. Ning suggested this pattern fits training on genre-specific data, though Witness cannot identify the exact data used by Anthropic. Source: thedecoder