Google DeepMind has introduced Gemini Robotics 2, its most advanced vision-language-action (VLA) model to date, designed to control robots of various shapes, including tabletop arms and humanoid forms. The model integrates image recognition, language processing, and action control to enable robots to operate in physical environments. Developers can apply for early access through a waitlist, according to the report. The company describes Gemini Robotics 2 as an 'intelligence layer' for adaptive robots, allowing them to manage full-body movement, perform fine motor tasks, and coordinate multiple units. The model is part of a broader effort to enhance robotic capabilities through advanced AI integration.

In addition to Gemini Robotics 2, DeepMind has also released Gemini Robotics ER 2, a model focused on embodied reasoning. This system enables robots to understand the physical world and make decisions based on that knowledge. ER 2 replaces the earlier ER 1.6 version, which was released in April. The new model is now available in Google AI Studio, allowing developers to experiment with its capabilities.

The report highlights DeepMind's ongoing commitment to advancing robotic systems through AI-driven innovation. According to the source, the models are intended to support a wide range of robotic applications, from simple mechanical arms to complex humanoid forms.

Source: thedecoder