RC RANDOM CHAOS

Google DeepMind's Gemini Robotics 2 adds whole-body control and robot teamwork

· via Hacker News

Original source

Gemini Robotics 2 brings whole body intelligence to robots

Hacker News →

Google DeepMind has released Gemini Robotics 2, a suite of three models meant to serve as the intelligence layer for adaptable robots. Where the prior generation handled upper-body, table-top manipulation, the new lineup extends to full humanoid control — walking, crouching, and balancing — plus finer dexterity across both five-fingered hands and simple two-finger grippers. The release spans a flagship vision-language-action (VLA) model that turns visual and language input into motor commands, an embodied-reasoning model (Gemini Robotics ER 2) that acts as a high-level planner, and an efficient on-device VLA optimized to run locally without network connectivity.

The headline capabilities are longer autonomous task sequences and multi-robot collaboration. The ER 2 reasoning model can plan and track multi-step jobs lasting several minutes and involving hundreds of decisions, recognize when tasks start and finish, self-correct after a failed step, and coordinate multiple robots of different types on a single workflow. The on-device model inherits DeepMind’s ‘motion transfer’ technique and can reportedly adapt to an entirely new robot body — with different shapes, sensors, and degrees of freedom — in a few hours using fewer than 200 examples. ER 2 is available now in Google AI Studio and in private preview on the Gemini Enterprise Agent Platform; the VLA and on-device models are limited to early-access partners.

DeepMind pairs the launch with a safety push, introducing ASIMOV-Agentic, a benchmark for agentic safety that measures whether the reasoning agent will refuse unsafe tool calls from the VLA, judge whether a task is even possible, and ask for human help when uncertain. The company says ER 2 is its strongest model yet at following safety constraints and detecting nearby humans to trigger a safe stop. The framing is explicit: DeepMind positions this as a step toward general-purpose ‘physical AI’ and AGI in the real world, though it concedes movement speed and human-level dexterity remain works in progress.

Read the full article

Continue reading at Hacker News →

This is an AI-generated summary. Read the original for the full story.