Google DeepMind’s new AI model can control a robot’s entire body

Google DeepMind’s new AI model can control a robot’s entire body

Google DeepMind 的全新 AI 模型可操控机器人全身

Google DeepMind says the latest version of its Gemini Robotics AI model can “control entire humanoid robots.” While the previous model focused on controlling a humanoid robot’s upper body, Gemini Robotics 2 now supports “whole-body motions” ranging from its feet to fingertips, according to an announcement on Thursday.

Google DeepMind 表示,其 Gemini Robotics AI 模型的最新版本能够“操控整个人形机器人”。据周四发布的消息,之前的模型主要侧重于控制人形机器人的上半身,而 Gemini Robotics 2 现在支持从脚尖到指尖的“全身动作”。

The new model will allow humanoid robots to perform a wider range of actions, as it allows them to walk, crouch, stretch, and manipulate objects. Videos shared by Google show how Apptronik’s Apollo 2 robot can bend over to pick up a watering can, as well as find and take specific items off a shelf.

该新模型将使人形机器人能够执行更广泛的动作,因为它允许机器人行走、蹲下、伸展并操纵物体。Google 分享的视频展示了 Apptronik 的 Apollo 2 机器人如何弯腰捡起喷壶,以及如何在货架上寻找并取下特定物品。

Though Google DeepMind notes that its robots “have more to advance in movement speed,” it adds that this update “is an important step towards the skills needed to complete more complex, real-world tasks that require whole-body coordination.”

尽管 Google DeepMind 指出其机器人在“移动速度方面还有待提升”,但它补充说,此次更新“是迈向完成需要全身协调的复杂现实任务所需技能的重要一步”。

Additionally, Gemini Robotics 2 supports better dexterity, as it can now control more complex, five-fingered hands. That enables robots to perform tasks like sealing a Ziploc, tying a trash bag, or unscrewing a lightbulb.

此外,Gemini Robotics 2 支持更好的灵活性,因为它现在可以控制更复杂的五指手部。这使得机器人能够执行诸如密封保鲜袋、系垃圾袋或拧下灯泡等任务。

Google DeepMind is updating Gemini Robotics ER (embodied reasoning) as well, a vision-language model that helps robots to analyze their surroundings, process instructions, and perform multi-step tasks. Gemini Robotics ER 2 is better at completing tasks over an extended period of time and “now understands when tasks begin and end.”

Google DeepMind 还在更新 Gemini Robotics ER(具身推理)模型,这是一种视觉语言模型,旨在帮助机器人分析周围环境、处理指令并执行多步骤任务。Gemini Robotics ER 2 在长时间完成任务方面表现更佳,并且“现在能够理解任务何时开始和结束”。

Google DeepMind says this update also allows multiple robots of different types to work together and complete tasks, with one video showing how Apollo 2 instructs Google’s dual-arm robot to put tools inside a bin while cleaning the garage.

Google DeepMind 表示,此次更新还允许不同类型的多个机器人协同工作并完成任务。其中一段视频展示了 Apollo 2 如何在清理车库时指挥 Google 的双臂机器人将工具放入箱子中。

The company notes Gemini Robotics ER 2 is its “safest robotics model to date,” as it can “better detect when humans are nearby, trigger safety tool calls and bring the robot to a safe stop if someone approaches too closely.”

该公司指出,Gemini Robotics ER 2 是其“迄今为止最安全的机器人模型”,因为它能“更好地检测到附近是否有人员,触发安全工具调用,并在有人过于靠近时使机器人安全停止”。

Meanwhile, Google DeepMind has brought improvements to its Gemini Robotics On-Device Model, which can run locally on a robot without an internet connection. This model can now adapt to new embodiments faster, including those with “drastically different shapes, sensors and degrees of freedom.”

与此同时,Google DeepMind 还对其 Gemini Robotics 端侧模型进行了改进,该模型无需互联网连接即可在机器人本地运行。该模型现在能更快地适应新的机体,包括那些具有“截然不同的形状、传感器和自由度”的机体。