Stanford's Terminal-Bench-Science tests AI agents on real research — top model hits 30%
Original source
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
Hacker News →Terminal-Bench-Science is a new benchmark from Stanford researchers and the Terminal-Bench team that measures how well AI agents handle actual scientific work rather than textbook problems. Its defining premise is that practicing scientists—not model vendors or data companies—should define what capable looks like. The first release ships 70 expert-curated tasks spanning the life, physical, Earth, mathematical, and engineering sciences, covering everything from statistical inference and simulation to theorem proving, signal processing, and scientific machine learning. Agents are dropped into realistic environments and graded on concrete artifacts—analyses, proofs, code, data products—via reproducible, task-specific tests.
The difficulty is deliberate and the results are sobering. The strongest system, Claude Opus 5 running Claude Code, resolves just 30% of tasks, trailed by GPT-5.6 Sol with Codex at 22.4% and Claude Fable 5 at 21.4%; most other models fall below 10%, with GLM 5.3 leading open models at 8.1%. Tasks were calibrated against current frontier models, pushing resolution rates more than 10 points below Terminal-Bench 3.0 for every model tested on both. The team also reports cost and token efficiency frontiers, noting that cheaper models sometimes match pricier ones—GPT-5.6 Sol matches Fable 5’s score at under a third of the cost.
The project is run as an open, continuous effort rather than a one-off paper. Tasks are proposed and reviewed on GitHub and Discord through a multi-stage gauntlet—domain reviewers, technical reviewers, and a final bar-raiser—that winnowed 920 proposals down to the 70 that shipped. Future releases will add and retire tasks, broaden domain coverage, and recalibrate against whatever frontier models exist at the time, with the stated goal of building a lasting feedback loop between scientific needs and AI development. Work on version 0.2 is already underway.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.