NewsAgg

local preview
technology

Google Introduces Gemini Robotics 2 for Intelligent Whole-Body Robot Control

New AI models enable robots to reason through complex tasks, perform dexterous manipulation, and collaborate with other robots without constant human oversight.

Google Introduces Gemini Robotics 2 for Intelligent Whole-Body Robot Control

Google has announced Gemini Robotics 2, an advancement in AI-powered robotic intelligence that expands beyond previous capabilities to enable whole-body control, advanced dexterity, and multi-robot teamwork.

Google Introduces Gemini Robotics 2 for Intelligent Whole-Body Robot Control

According to Google’s announcement, most existing robots are pre-programmed or teleoperated for narrow, repetitive tasks and lack the ability to truly learn or adapt to unpredictable environments. Gemini Robotics 2 addresses this limitation by providing robots with AI models that enable them to think, act, and interact intelligently.

The system comprises three models. Gemini Robotics 2 is a vision-language-action model (VLA) that converts vision and language input into motor control, enabling robots to control full humanoid bodies from feet to fingertips. Gemini Robotics ER 2 is an embodied reasoning model that acts as a high-level decision-maker, allowing robots to understand the physical world, plan multi-step tasks, and collaborate with other robots. Gemini Robotics On-Device 2 is optimized to run locally on robotic devices without requiring network connectivity.

In demonstrations, the system showed whole-body coordination tasks. For example, when controlling Apptronik’s Apollo 2 humanoid robot, the system could interpret an instruction like “put the watering can into the green bin in the bottom shelf,” then execute a sequence of movements including walking to a table, picking up the object, and placing it at the destination. The system can also control the 22 degree-of-freedom SharpaWave hand to perform delicate actions like tying knots or sealing ziplock bags.

A key advancement is the ability to execute longer task sequences lasting several minutes and involving hundreds of decisions. According to Google, Gemini Robotics ER 2 can now understand when tasks begin and end and identify key events, representing a step change in progress understanding. The system also enables different types of robots to communicate and work together on complex workflows.

For on-device deployment, Gemini Robotics On-Device 2 can adapt to new bi-arm robot embodiments with just a few hours of adaptation time, typically using fewer than 200 examples. This adaptation works even with robots that have drastically different shapes, sensors, and degrees of freedom.

Google has introduced ASIMOV-Agentic, a new safety benchmark for agentic systems, measuring the embodied reasoning agent’s ability to refuse unsafe actions and predict task feasibility. According to the announcement, Gemini Robotics ER 2 showed improvements in safety constraint following and human proximity detection, including the ability to trigger safety stops when humans approach too closely.

Key facts

  • Gemini Robotics 2 enables robots to control entire humanoid bodies from feet to fingertips, expanding beyond previous upper-body-only capabilities
  • The system comprises three models: a vision-language-action model for motor control, an embodied reasoning model for task planning, and an on-device efficient model
  • Robots can now execute complex multi-step tasks lasting several minutes with hundreds of decisions
  • Gemini Robotics On-Device 2 can adapt to new robot embodiments in just a few hours using fewer than 200 examples
  • The system introduces multi-robot collaboration, allowing different types of robots to work together on complex workflows
  • Safety improvements include the ability to detect nearby humans and trigger safety stops when someone approaches too closely

Sources

← All posts