ARC-AGI-3 Leaderboard Grades AI Agents on Efficiency, Not Just Raw Scores
ARC Prize has published its ARC-AGI-3 leaderboard, marking a shift in what the benchmark measures. Where ARC-AGI-1 and 2 tested passive “fluid intelligence” on static puzzles, version 3 drops AI agents into novel interactive environments and scores how well they adapt on the fly. The framing argument is that genuine intelligence is defined not only by solving hard problems but by solving them cheaply, so the leaderboard plots performance against cost-per-task rather than accuracy alone.
Entries are split into three classes. Reasoning systems appear as connected points that trace a single model across different thinking budgets, illustrating the diminishing returns as reasoning time climbs toward an asymptote. Base LLMs — such as GPT-4.5 and Claude 3.7 — are shown as single-shot inference with no extended reasoning, representing raw model capability. Kaggle systems are purpose-built competition submissions constrained to a $50 compute budget across 120 evaluation tasks, standing in as the efficiency-first end of the field.
Several caveats limit how the numbers should be read. Only systems costing under $10,000 to run are displayed, and any model that failed to produce complete outputs had its remaining tasks scored as incorrect. Results labeled “preview” are unofficial and may rest on partial testing, some ARC-AGI-2 figures are estimates derived from limited runs and o1-pro pricing, and Gemini 3 Pro’s costs are provisional placeholders pending a retest once the model ships.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.