DeepMind has launched Gemini Robotics ER 2, an advanced model designed to enhance robots' ability to perform tasks in the physical world. This model enables robots to understand video input, plan multi-step tasks, and collaborate with other robots. It also allows for real-time decision-making and integration with tools like Google Search. The release marks a significant upgrade from the previous version, Gemini Robotics ER 1.6, by enabling continuous video tracking and adaptive task execution.
Gemini Robotics ER 2 improves tool orchestration workflows, outperforming ER 1.6 in three control modes: real VLA, sim VLA, and human tele-op. The model integrates with the Gemini Live API, using a bidirectional streaming endpoint optimized for latency-sensitive tasks. This results in fluid orchestration, allowing the model to command action models and robotics APIs to complete multi-step tasks without the jarring 'stop-and-think' pauses. A demo with Spot from Boston Dynamics showcases the model's ability to fetch objects based on natural language commands.
The release includes enhancements in spatial reasoning, progress classification, and moment-finding. Gemini Robotics ER 2 achieves 57.4% accuracy on progress classification tasks and 91.3% accuracy on moment-finding tasks, with a 0.96s mean absolute distance. It also supports multi-robot collaboration, enabling diverse machines to communicate and complete complex tasks together. The model is now available to developers via the Gemini API, Google AI Studio, and the Gemini Enterprise Agent Platform.
Source: deepmind