Gemini Robotics ER 2 revolutionises robot capabilities
Gemini Robotics ER 2 marks a significant advancement in how robots can understand videos, organise tasks, and collaborate with multiple robots, enhancing their utility in real-world scenarios.
Google recently introduced Gemini Robotics ER 2, a new model that serves as a sophisticated control system for robots. This model facilitates rapid spatial reasoning, detailed multi-step task planning, and enables cooperation among various robots. Developers can access this model through the Gemini API, Google AI Studio, or the Gemini Enterprise Agent Platform to create effective physical AI agents.
The Gemini Robotics ER 2 model enables robots to monitor video feeds, track progress, and correct errors instantly. This innovation aims to make robots more secure and beneficial in everyday environments. Real-time decision-making is crucial for robots to function effectively alongside humans, and this model addresses that need by offering enhanced reasoning capabilities.
The latest version, Gemini Robotics ER 2, builds on its predecessor by allowing robots to communicate with humans, understand their surroundings, and execute complex plans. It performs high-level reasoning and delegates motor functions to specialised vision-language-action (VLA) models. Additionally, it can access tools like Google Search or user-defined functions, enabling it to think ahead while executing tasks.
This model significantly improves over the earlier version, ER 1.6, by allowing robots to monitor continuous video feeds, adapt when issues arise, and efficiently transition between tasks. It also supports multi-robot collaboration, enabling coordinated efforts to complete intricate tasks that a single robot cannot manage alone.
Gemini Robotics ER 2 is currently available for developers through various platforms, including the Gemini API and Google AI Studio. To aid developers, Google provides examples of how to configure the model for more practical AI applications.
Enhancing task execution and collaboration
Gemini Robotics ER 2 is designed to handle complex, multi-step tasks by orchestrating the necessary steps and enabling robots to self-correct and adapt to new situations. Developers can employ low-level control interfaces, such as VLA models or navigation APIs, as tools, and stream varied media directly into the model for better orchestration.
This model demonstrates superior performance in tool orchestration across three modes: real VLA, simulated VLA, and human teleoperation. Its high-level reasoning is synchronised with execution speed, resulting in seamless task orchestration without interruptions.
Google has partnered with Boston Dynamics to demonstrate Gemini Robotics ER 2’s capabilities using the Spot robot, showcasing how it can fetch items on command using the model’s advanced orchestration capabilities.
Progress tracking and moment finding
One of the challenges in robotics is determining task completion. Gemini Robotics ER 2 introduces improvements in video comprehension and progress tracking, ensuring complex tasks meet specifications before moving on. The model excels in progress classification and moment finding, offering real-time situational awareness and allowing for on-the-fly adjustments.
The progress classification feature segments video feeds into progress levels, while the moment-finding capability allows precise task switching. Gemini Robotics ER 2 achieves impressive accuracies in these tasks, outperforming previous models and offering speed advantages.
Multi-robot collaboration and safety
Not all robots are suited for every task; however, Gemini Robotics ER 2 facilitates collaboration between different robotic systems, such as Apptronik’s Apollo 2 and Franka F3 Duo, by enabling communication through shared semantic understanding.
The model excels in spatial reasoning, achieving high benchmarks for success detection, question answering, and instrument reading. It prioritises safety, performing well on benchmarks related to safety instruction following and human proximity, ensuring safe operation in real-world environments.
As part of its commitment to safety, Google has developed a benchmark to evaluate foundational models’ abilities to enforce safety constraints and ensure physical feasibility. The goal is to push these models towards handling even more intricate tasks, advancing the development of useful robotic systems.
Stay updated with product news, event details, and special offers by subscribing to our newsletters. You can manage your subscription preferences at any time.
