Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration

Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration

Gemini Robotics ER 2:通过视频理解、任务编排与多机器人协作赋能机器人技术

Introducing Gemini Robotics ER 2. Gemini Robotics ER 2 represents a step change in powering robots with video understanding, task orchestration, and multi-robot collaboration — making it possible for robots to be more helpful in the physical world. 隆重推出 Gemini Robotics ER 2。Gemini Robotics ER 2 代表了在视频理解、任务编排和多机器人协作方面赋能机器人的重大飞跃,使机器人能够在物理世界中发挥更大的作用。

We are launching Gemini Robotics ER 2, a new model designed to act as a high-level brain for robots. It enables real-time spatial reasoning, multi-step task planning, and collaboration between different robots. You can access the model now via the Gemini API, Google AI Studio, or the Gemini Enterprise Agent Platform to start building your own physical AI agents. 我们正式发布 Gemini Robotics ER 2,这是一款旨在充当机器人“高级大脑”的新模型。它支持实时空间推理、多步骤任务规划以及不同机器人之间的协作。您现在可以通过 Gemini API、Google AI Studio 或 Gemini Enterprise Agent Platform 访问该模型,开始构建您自己的物理 AI 智能体。

For robots to assist humans in everyday environments, accurate spatial reasoning is not enough. Robots must also think fast, timing their decisions and reasoning with the real-time speed of the physical world. That’s why today we’re launching Gemini Robotics ER 2, our most capable “embodied reasoning” model for robotics. 为了让机器人能够在日常环境中协助人类,仅有精确的空间推理是不够的。机器人还必须能够快速思考,以物理世界的实时速度来安排决策和推理。因此,我们今天推出了 Gemini Robotics ER 2,这是我们目前针对机器人技术功能最强大的“具身推理”模型。

Think of Gemini Robotics ER 2 as a high-level brain for robots. It allows robots to chat with humans, understand the physical world, and plan multi-step tasks. It then hands off motor execution to any given lower-level vision-language-action (VLA) model. Gemini Robotics ER 2 can also natively call tools like Google Search to find information, or any other user-defined function. The design of Gemini Robotics ER 2 allows the robot to “think” about what comes next while simultaneously performing its actions. 可以将 Gemini Robotics ER 2 视为机器人的高级大脑。它允许机器人与人类交流、理解物理世界并规划多步骤任务,随后将电机执行任务移交给任何指定的底层视觉-语言-动作 (VLA) 模型。Gemini Robotics ER 2 还可以原生调用 Google 搜索等工具来查找信息,或调用任何其他用户定义的函数。Gemini Robotics ER 2 的设计使得机器人能够在执行动作的同时“思考”下一步该做什么。

Gemini Robotics ER 2 represents a significant upgrade over Gemini Robotics ER 1.6. By watching continuous video feeds, robots can now track their own progress, adapt if something goes wrong, and know exactly when to move on to the next step. We are also introducing multi-robot collaboration, enabling robots to work together in shared spaces and complete complex workflows a single robot could not do alone. Gemini Robotics ER 2 相比 Gemini Robotics ER 1.6 实现了重大升级。通过观察连续的视频流,机器人现在可以跟踪自己的进度,在出现问题时进行调整,并准确知道何时进入下一步。我们还引入了多机器人协作功能,使机器人能够在共享空间中协同工作,完成单台机器人无法独立完成的复杂工作流程。

Advancing physical agentic capabilities. Most tasks in the physical world are complex and require multiple steps to complete. Gemini Robotics ER 2 is a physical agent, orchestrating steps for the robot and enabling it to self-correct, and generalize to more novel situations. To build an agentic setup, developers can declare low-level control interfaces — like Vision-Language-Action (VLA) models or navigation APIs — as tools, and stream multimodal video, audio, or text directly into the model. 推进物理智能体能力。物理世界中的大多数任务都很复杂,需要多个步骤才能完成。Gemini Robotics ER 2 是一个物理智能体,它为机器人编排步骤,使其能够自我纠正并推广到更多新颖的情况。为了构建智能体设置,开发人员可以将底层控制接口(如 VLA 模型或导航 API)声明为工具,并将多模态视频、音频或文本直接流式传输到模型中。

In robotics, high-level reasoning depends on execution speed. Gemini Robotics ER 2 integrates into the Gemini Live API, using a bidirectional streaming endpoint optimized for latency-sensitive tasks. The result is fluid orchestration: Gemini Robotics ER 2 commands action models and robotics APIs to complete multi-step tasks without the jarring “stop-and-think” pauses. 在机器人技术中,高级推理依赖于执行速度。Gemini Robotics ER 2 集成了 Gemini Live API,使用针对延迟敏感型任务优化的双向流式端点。其结果是流畅的编排:Gemini Robotics ER 2 指挥动作模型和机器人 API 完成多步骤任务,而不会出现令人不适的“停顿思考”间隙。

Unlocking temporal intelligence for robust task completion. One of robotics’ hardest challenges is knowing when a task is done. Gemini Robotics ER 2 brings a step-change in video understanding and progress tracking to verify that complex tasks — such as tightening a light bulb or tying a trash bag — are complete to specification before switching to the next task. 解锁时间智能以实现稳健的任务完成。机器人技术中最困难的挑战之一是判断任务何时完成。Gemini Robotics ER 2 在视频理解和进度跟踪方面带来了质的飞跃,能够验证复杂任务(例如拧紧灯泡或系好垃圾袋)是否按规范完成,然后再切换到下一个任务。

Continuous progress classification. Progress classification refers to a robot’s ability to track progress towards task completion. In our evaluations, we assign each frame in a video feed into five levels of progress (0-20%, 20-40%, 40-60%, 60-80%, 80-100%). By quantifying task progress, Gemini Robotics ER 2 provides robots with real-time situational awareness, and allows them to adjust actions on the fly or retry failed steps without restarting an entire workflow. 持续进度分类。进度分类是指机器人跟踪任务完成进度的能力。在我们的评估中,我们将视频流中的每一帧分配到五个进度级别(0-20%、20-40%、40-60%、60-80%、80-100%)。通过量化任务进度,Gemini Robotics ER 2 为机器人提供了实时态势感知,并允许它们动态调整动作或重试失败的步骤,而无需重新启动整个工作流程。

Precision moment-finding. Moment-finding measures a model’s ability to identify the exact video frame where a critical event takes place (i.e. when to stop pouring coffee into a cup). Gemini Robotics ER 2 achieves significant gains in performance on moment finding, enabling robots to precisely switch between tasks, verify success and suggest corrections. 精确时刻定位。时刻定位衡量的是模型识别关键事件发生的确切视频帧的能力(例如,何时停止向杯中倒咖啡)。Gemini Robotics ER 2 在时刻定位方面取得了显著的性能提升,使机器人能够精确地在任务之间切换、验证成功并提出修正建议。