GPT-6 Astra tops ARC-AGI-3, beating human action efficiency on 96% of levels
OpenAI’s GPT-6 Astra posted state-of-the-art results on ARC-AGI-3, the ARC Prize benchmark for agentic intelligence that drops models into novel, turn-based environments and forces them to explore, infer the goal, and build a working model with no instructions. Running under ARC Prize’s Standard harness, Astra scored 62.7% for roughly $26K; under the Provider Adapter harness, which preserves opaque reasoning state across requests, it reached 99.9% for about $19K. Humans solve 100% of these environments, and the paid-tester baseline works out to about $12.78 per game attempted — though ARC Prize notes that if you price only the brain’s energy as electricity, the human cost drops to a fraction of a cent.
The headline finding is efficiency, not just completion. Against a baseline drawn from roughly 500 members of the public, Astra used fewer actions than the median human on 96% of levels and averaged 51.7% fewer actions per level — crossing a line ARC Prize had bet would separate humans from AI. The team’s earlier hypothesis was that even a model that solved an environment would need far more exploration than a person; instead, frontier AI showed a near-binary pattern: once it grasps the mechanics, it executes within human range. Replays show why, with Astra distilling each scene into a dense, code-like algebraic notation — tracking objects, coordinates, rotation indices, and ordered multi-step plans in a shorthand it invents on the fly.
In a separate red-teaming harness called PRO-LONG, where Astra had a code sandbox, it went further and wrote game-specific software — maze solvers, combat and patrol models, and state-checking scripts — occasionally assembling small custom libraries per game. ARC Prize flags that those runs aren’t comparable to the human tests, since participants had no code interpreter or scratchpad. Even setting the tool-assisted results aside, the action-efficiency milestone is the notable one: it narrows the “residual gap” ARC-AGI is designed to measure, suggesting frontier models are closing in on human-like sample efficiency at learning unfamiliar environments, not just raw task success.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.