Details
- Google DeepMind announced Gemini Robotics 2, a next-generation physical AI system positioned as a single "brain" that can run across different robots.
- The launch introduces three models: Gemini Robotics 2, Gemini Robotics ER 2, and Gemini Robotics On-Device 2, each covering different layers of perception, reasoning, and motor control.
- Gemini Robotics 2 is a vision-language-action model that controls full humanoids "from feet to fingertips," enabling whole-body motions such as walking, crouching, stretching and object manipulation beyond tabletop tasks.
- Gemini Robotics ER 2 is an embodied reasoning model focused on real-world video understanding and multi-step planning, acting as a high-level agent that can track progress and adjust tasks over several minutes.
- Gemini Robotics On-Device 2 is optimized to run locally on robotic hardware, adapting to new robot bodies within hours while enabling low-latency control and privacy-preserving, on-device intelligence.
- DeepMind highlights advanced dexterity, showing Gemini Robotics 2 controlling five-fingered hands to tie knots or screw lightbulbs, while also managing simpler parallel grippers for complex manipulation workflows.
- In partnership demonstrations, Apptronik's Apollo 2 humanoid uses a single natural-language prompt to coordinate whole-body movement — reaching, bending, and picking up a watering can — illustrating general-purpose, prompt-driven physical intelligence.
- The new models are designed to support multi-robot collaboration, allowing robots to work together as a team on coordinated tasks in homes, workplaces, and industrial settings.
- Gemini Robotics 2 builds on Gemini 2.0 multimodal capabilities by adding robotic control as an output, translating language and visual understanding directly into low-level motor commands for diverse embodiments.
Impact
By bringing whole-body intelligence, dexterous manipulation, and on-device adaptation into a unified Gemini Robotics 2 stack, Google DeepMind strengthens Alphabet’s position in embodied AI against rivals like NVIDIA, Figure, and Tesla’s Optimus efforts. Integrated embodied reasoning and local control could accelerate real-world deployment of humanoids in logistics, manufacturing, and service environments, while showcasing how frontier foundation models are evolving from digital assistants into physical agents. The multi-model architecture also positions Gemini Robotics as a platform layer for robot OEMs, potentially reshaping how robotics vendors design control software around large multimodal models.