AI benchmark leaderboard adds private tests to curb gaming; Claude Fable 5.1 leads
Artificial Analysis has pushed out v4.2 of its Intelligence Index, an interim refresh meant to keep the benchmark relevant while the frontier moves faster than its planned v5 cycle. The update pulls in two harder, more realistic evaluations: AA-Briefcase, an in-house test of agentic knowledge work built by industry experts around multi-week projects with thousands of linked source files, and Surge AI’s GDP.pdf, which forces models to reason across 4,592 pages of professional documents graded against 1,275 expert-written criteria. GPQA Diamond was dropped after scores saturated it.
The headline methodology change is defensive: 40 percent of the Index weighting now comes from private, held-out test sets, double the share in v4.1, specifically to make it harder for labs to tune models to public benchmarks. Artificial Analysis also reworked its grading pipeline, adding system prompts, fixing answer-key errors, re-anchoring Elo scales for stability, and hardening code-execution sandboxes so slow-but-correct solutions no longer register as failures.
On results, Anthropic’s Claude Fable 5.1 tops the overall Index, with OpenAI’s GPT-6 Astra second on a four-point gain over its predecessor, followed by Meta, SpaceXAI, Moonshot/Kimi, Z.AI, and Google. GPT-6 Astra stands out for token efficiency near the intelligence frontier, while Anthropic leads the new agentic AA-Briefcase eval and OpenAI leads GDP.pdf at 33.2 percent. The bigger signal is the industry’s shift toward private, agentic, long-context tests as public leaderboards lose their power to distinguish frontier models.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.