Google has introduced Gemini Robotics 2.0, an upgraded version of its AI-powered robotics models that promises improved dexterity and safety. The new release enables robots to perform more complex tasks, continuously analyze their environments, and collaborate with other machines. This advancement comes with a trio of new sub-models, one of which is now publicly available for developers. The release marks a step closer to creating generalist robots capable of executing a wide range of tasks, as envisioned by Google DeepMind scientists.
The core of the upgrade is Gemini Robotics ER 2, an enhanced 'embodied reasoning' model that integrates with the Gemini Live API. This model is designed to understand instructions and the surrounding environment, processing live video feeds from the robot's cameras. According to Google, ER 2 can classify video frame completeness with almost 60 percent accuracy, a marked improvement over the previous 1.6 release and competing models. The model also excels at identifying key moments in video feeds, achieving nearly 90 percent accuracy in tasks like pouring a cup of coffee, allowing robots to adjust their actions in real time.
The development of Gemini Robotics 2.0 is part of Google's broader effort to create physical AI systems that can operate safely alongside humans. The company emphasizes the inclusion of traditional physical safety measures alongside AI safety frameworks, with a new safety benchmark called ASIMOV-Agentic. This benchmark evaluates models across various safety factors, including the ability to refuse unsafe tool calls and request human assistance when uncertain. Google DeepMind claims that Gemini Robotics ER 2 is its safest model yet, demonstrating a strong ability to understand safety and halt actions when a human is too close to the robot.
Source: arstechnica