NewsAgg

local preview
technology

Google's Gemini Robotics 2 enables whole-body control and multi-robot teamwork

New AI models give robots advanced dexterity, reasoning, and coordination abilities to complete complex multi-step tasks in unpredictable environments.

Google's Gemini Robotics 2 enables whole-body control and multi-robot teamwork

Google has introduced Gemini Robotics 2, a suite of AI models designed to enable robots with intelligent whole-body control, advanced dexterity, and multi-robot collaboration capabilities.

Google's Gemini Robotics 2 enables whole-body control and multi-robot teamwork

According to Google DeepMind, most robots today are pre-programmed or teleoperated for narrow, repetitive tasks and lack the ability to truly learn or adapt to unpredictable environments. Gemini Robotics 2 aims to address this limitation by providing an intelligence layer that allows robots of different shapes and sizes to think, act, and interact intelligently.

The system consists of three models. Gemini Robotics 2 is the core vision-language-action model that converts vision and language input into motor control, enabling full humanoid control from feet to fingertips, as well as dexterous manipulation on hands and grippers. Gemini Robotics ER 2 is an embodied reasoning model that acts as a high-level brain, enabling robots to communicate with humans, understand the physical world, and plan multi-step tasks lasting several minutes. Gemini Robotics On-Device 2 is an efficient model optimized to run locally on robotic devices without network connectivity.

In demonstrations, the system enabled Apptronik’s Apollo 2 humanoid robot to process instructions like “put the watering can into the green bin on the bottom shelf,” then execute the task by walking to a table, picking up the watering can, and placing it precisely at the destination. The models can also control fine-fingered hands to perform delicate actions such as tying knots or sealing ziplock bags, and operate grippers to perform complex packing tasks.

A significant advancement is multi-robot collaboration, which allows different robot types to communicate and work together on complex workflows that single robots cannot complete alone. The embodied reasoning model now understands when tasks begin and end, and can execute longer task sequences lasting several minutes involving hundreds of decisions.

For on-device adaptation, Gemini Robotics 2 can adapt to new robot embodiments in just a few hours using fewer than 200 examples, even when robots have drastically different shapes, sensors, and degrees of freedom.

Google DeepMind emphasized that safety is foundational to the robotics research. The company introduced ASIMOV-Agentic, a new benchmark for agentic safety orchestration, and noted that Gemini Robotics ER 2 can detect nearby humans and trigger safety stops if someone approaches too closely.

Key facts

  • Gemini Robotics 2 enables whole-body control of humanoid robots, allowing tasks like walking, crouching, and object manipulation
  • The system includes three models: a vision-language-action model for control, an embodied reasoning model for planning, and an on-device efficient model
  • Robots can now execute multi-step tasks lasting several minutes and collaborate with other robots to solve complex workflows
  • The on-device model can adapt to new robot embodiments in hours using fewer than 200 examples
  • Gemini Robotics ER 2 can detect nearby humans and bring robots to safe stops when people approach closely

Sources

← All posts