Details
- Google DeepMind launched Gemini Robotics ER 2, a new embodied-reasoning model for robots that acts as a high-level planner for multi-step physical tasks.
- The model is designed to handle video understanding, spatial reasoning, task orchestration, and multi-robot collaboration while handing motor execution to lower-level vision-language-action models.
- Google says ER 2 can stream multimodal inputs, call tools such as Google Search or developer-defined functions, and integrate through the Gemini API, Google AI Studio, and the Gemini Enterprise Agent Platform.
- Compared with Gemini Robotics ER 1.6, ER 2 adds continuous video-based progress tracking, improved moment finding, and better tool orchestration across real robots, simulations, and teleoperation.
- Google reports 57.4% accuracy on progress classification and 91.3% accuracy on moment finding, and says ER 2 runs at roughly 4x the execution speed of much larger models for latency-sensitive robotics use.
Impact
ER 2 pushes robotics toward a more agentic stack: fast reasoning at the top, specialized control below, and tighter closed-loop self-correction in between. If the reported latency and progress-tracking gains hold up outside demos, it could accelerate adoption of physical AI agents in warehouses, labs, and service robotics over the next 12–24 months.