Anthropic's Claude Opus 5 achieved 30.2% on ARC-AGI-3, the reasoning benchmark designed to measure fluid intelligence on novel tasks, nearly quadrupling OpenAI's GPT-5.6 Sol's previous record of 7.8%. The score was independently verified by the ARC Prize Foundation. ARC-AGI-3 presents tasks models have never encountered in training—each environment is guaranteed novel, making memorization irrelevant and forcing genuine rule inference from first principles.
During testing, Opus 5 displayed reasoning behavior not previously documented in frontier models: it independently formulated algebraic reflection equations, including the equation 4_center = 2×axis − 5_center without external prompting. The model solved five previously unsolved environments, four of them at or above human efficiency, bringing the total number of solved public-demo environments from one to six. On older benchmarks, Opus 5 scores 97.5% on ARC-AGI-1 and 90.4% on ARC-AGI-2.
For production AI architects, Opus 5's breakthrough on abstract reasoning and novel-problem-solving signals measurable progress on a benchmark specifically designed to resist benchmark-gaming techniques. The ARC-AGI suite filters out memorization and pattern-matching by construction; a 30% jump over the previous frontier represents gains in architectural reasoning capacity and inference-time search strategy, not just training-data volume.