RC RANDOM CHAOS

GPT-6 Astra dominates block-in-bowl robotics test, but stalls on the puzzle like Fable

· via Hacker News

Original source

GPT-6 Astra on robot arms

Hacker News →

OpenAI’s GPT-6 Astra took control of the same YAM robot arms used in an earlier Claude Fable comparison, running under the Inspect Robots agent policy on two manipulation tasks. On the simpler job — placing a red block into a bowl — Astra succeeded 19 times out of 20, far outpacing Fable 5.1 (8/20) and Fable 5 (1/20). It was also faster and cheaper, averaging 2.5 minutes and roughly $0.94 per run versus Fable 5.1’s 6.8 minutes and $2.12.

The harder task exposed a shared ceiling. Inserting a knobbed blue puzzle piece into a matching groove tripped up both models almost equally: Astra completed it 2 times in 20, the same as Fable 5.1, with both systems reaching the groove and then stalling at the identical final step. Fine, contact-rich insertion remains an open problem regardless of which frontier model is driving. Human graders scored every trial by the highest stage it reached, so even failed runs recorded partial progress across the 120 total trials.

The authors are candid about the caveats. Runs weren’t interleaved, the bowl comparison used a different rig for the Fable trials, and grading was operator-judged with the model identity known, leaving room for unconscious bias. Costs are list price with no prompt caching applied to the Anthropic runs, while OpenAI auto-cached about a fifth of Astra’s input without discount — meaning Astra’s cost advantage is, if anything, understated. All models ran at medium reasoning effort only, so the results are a snapshot rather than a ceiling.

Read the full article

Continue reading at Hacker News →

This is an AI-generated summary. Read the original for the full story.