AI breakthroughs in robotics won’t change your life any time soon

本文为原文前 6,000 字符的节选翻译,完整内容请查看原文。

The story is a collaboration between MIT Technology Review and Aventine, a non-profit research foundation that creates and supports content about how technology and science are changing the way we live.

这篇报道是由《麻省理工科技评论》与非营利研究基金会 Aventine 合作完成的,该基金会致力于创作并支持有关科技如何改变人类生活方式的内容。

A robot shaped like a human—white with a black head and torso—has been popping up on video feeds. Perhaps you’ve seen it dance or pass popcorn, put trash in a bin, vacuum, or press the button of a microwave. Or maybe you’ve watched it fall backward while handing out water bottles or struggle to iron a shirt.

一个外形像人、有着白色身体以及黑色头部和躯干的机器人频繁出现在视频流中。或许你曾见过它跳舞、递爆米花、把垃圾扔进桶里、吸尘或按下微波炉按钮。又或许,你曾看过它在分发水瓶时向后摔倒,或是笨拙地熨烫衬衫。

This would be Tesla’s Optimus, an AI-powered humanoid robot that Elon Musk, the company’s CEO, believes will be “not just Tesla’s biggest product ever, but probably the biggest product ever,” headed to work on factory floors and, later, in our homes. Eventually it “will have human and then superhuman dexterity,” he told shareholders in July.

这就是特斯拉的“擎天柱”(Optimus),一款由人工智能驱动的人形机器人。特斯拉首席执行官埃隆·马斯克认为,它“不仅将是特斯拉有史以来最大的产品,甚至可能是史上最大的产品”,未来将投入工厂车间,随后进入我们的家庭。他在七月份告诉股东,它最终“将拥有人类乃至超越人类的灵巧度”。

Optimus robots could automate almost all human labor—from hauling sheet metal to folding laundry—for as little as $20,000 each, Musk argues. Speaking at the World Economic Forum’s annual meeting in Davos, Switzerland, in January, he predicted they could be on sale to the public by the end of 2027.

马斯克认为,擎天柱机器人可以实现几乎所有人类劳动的自动化——从搬运金属板到折叠衣物——每台成本仅需 2 万美元。今年一月,他在瑞士达沃斯举行的世界经济论坛年会上发言时预测,这些机器人最早可能在 2027 年底向公众发售。

Musk is not alone in his evangelism. Marc Andreessen, cofounder and general partner of the Silicon Valley venture capital firm Andreessen Horowitz, has said that robotics could become the “biggest industry in the history of the planet.” In January, Jensen Huang, CEO of Nvidia, said that humanoid robots would match human-level ability this year.

马斯克并非唯一的热衷者。硅谷风险投资公司 Andreessen Horowitz 的联合创始人兼普通合伙人马克·安德森曾表示,机器人技术可能成为“地球历史上最大的产业”。今年一月,英伟达首席执行官黄仁勋也表示,人形机器人将在今年达到人类水平的能力。

According to Morgan Stanley, the number of robots that “resemble and act like humans” is likely to reach nearly 1 billion by 2050, creating a market worth over $5 trillion.

据摩根士丹利预测,到 2050 年,“外形和行为像人”的机器人数量可能会达到近 10 亿台,创造一个价值超过 5 万亿美元的市场。

Such proclamations are in large part fueled by the idea that the same AI advances behind tools like OpenAI’s ChatGPT and Anthropic’s Claude will enable a new generation of robots to imitate human movement the way chatbots imitate human language.

这些声明在很大程度上源于一种观点:即推动 OpenAI 的 ChatGPT 和 Anthropic 的 Claude 等工具背后的 AI 进步,将使新一代机器人能够像聊天机器人模仿人类语言那样,模仿人类的动作。

But many robotics researchers are skeptical, arguing that such assumptions minimize the challenges of using an intelligence built on language and images to master the infinite variability of the physical world.

但许多机器人研究人员对此持怀疑态度,认为这种假设低估了利用基于语言和图像构建的智能来掌握物理世界无限变数所面临的挑战。

“None of those companies [building humanoid robots]—absolutely none of them—has any idea how to make those robots smart enough to be useful,” Yann LeCun, often referred to as one of the godfathers of AI, said at another event during the January Davos conference.

“那些(制造人形机器人的)公司——绝对没有一家——知道如何让这些机器人变得足够聪明以发挥实际用途,”常被称为 AI 教父之一的杨立昆在今年一月达沃斯会议期间的另一场活动上说道。

Researchers also point out that the tendency to conflate humanoid robots made to resemble people with so-called generalist machines able to learn and perform multiple tasks is misleading.

研究人员还指出,将外形像人的机器人与能够学习并执行多种任务的所谓“通用机器”混为一谈,这种倾向具有误导性。

”It’s very easy to make a robot that looks like a person,” explains Jonathan Hurst, cofounder and chief robot officer of Agility Robotics and professor of robotics at Oregon State University. “It is dramatically more difficult to make a machine that moves or behaves dynamically or physically like a person.”

“制造一个看起来像人的机器人非常容易,”Agility Robotics 的联合创始人兼首席机器人官、俄勒冈州立大学机器人学教授乔纳森·赫斯特解释道,“但要制造一台在动态或物理行为上像人一样运动的机器,难度要大得多。”

These tensions—over whether all-purpose humanoid robots are just around the corner or nowhere in sight, and whether current forms of AI are all that’s needed to perfect them—are playing out in robotics labs across the country, where the hype over timelines is obscuring painstaking but meaningful progress.

这些矛盾——关于全能型人形机器人是近在咫尺还是遥不可及,以及当前的 AI 形式是否足以完善它们——正在全国各地的机器人实验室中上演,而围绕时间表的炒作掩盖了那些艰苦但有意义的进展。

A decade or so ago, a series of breakthroughs led to a generative AI revolution that turned the long-imagined possibility of artificial intelligence into reality. Roboticists—though they disagree on exactly when this will happen—believe that an equally transformative revolution is possible in robotics, one that will endow machines with physical intuition and fluidity that has long been out of reach.

大约十年前,一系列突破引发了生成式 AI 革命,将长期以来对人工智能的构想变为了现实。机器人专家们——尽管他们对具体实现时间存在分歧——相信机器人领域也可能发生同样具有变革意义的革命,赋予机器长期以来难以企及的物理直觉和流畅性。

As progress in robotics inches forward, the question is whether the same methods and tools that fueled advances in AI are enough to get there, or if an entirely new path is required.

随着机器人技术的缓慢推进,问题在于推动 AI 进步的相同方法和工具是否足以实现这一目标,还是需要一条全新的路径。

To see one of the smartest robot brains working today, it’s worth looking at what Google DeepMind can do with a piece of equipment called ALOHA 2, short for “A Low-cost Open-source Hardware System for Bimanual Teleoperation.”

要了解当今最智能的机器人大脑之一,值得看看 Google DeepMind 利用一种名为 ALOHA 2 的设备所取得的成果,它是“用于双臂远程操作的低成本开源硬件系统”的缩写。

Roboticists have long clashed over whether a humanlike form is necessary for generalist robots, with proponents arguing that it will help them slot into the world as it exists and detractors saying it’s not worth the trouble. ALOHA 2 reflects this second way of thinking. Not much to look at, it’s just a pair of arms, some grippers, and a couple of cameras.

机器人专家长期以来一直在争论通用机器人是否必须具备人形,支持者认为这有助于它们融入现有的世界,而反对者则认为这不值得费力。ALOHA 2 反映了后一种思路。它看起来并不起眼,只是一对机械臂、一些抓手和几台摄像头。

But despite its seeming simplicity, it is a workhorse for researchers at Google DeepMind, who use it to test their most advanced AI for robotics system, Gemini Robotics, in their various labs.

尽管看起来简单,但它却是 Google DeepMind 研究人员的得力工具,他们在各个实验室中使用它来测试其最先进的机器人 AI 系统——Gemini Robotics。

When controlled by Gemini Robotics, ALOHA 2 becomes more of a generalist robot, in the sense that it can perform any number of tasks based on examples it’s been trained on. Ask it to pack a lunchbox and, as evidenced by a video of this exercise, it can use two pincer grippers to delicately place a piece of white bread into a Ziploc bag, close it, place a bunch of grapes in a Tupperware container, secure the lid, and then carefully move the items into a lunchbox before zipping it up.

在 Gemini Robotics 的控制下,ALOHA 2 变得更像是一个通用机器人,因为它能够根据训练示例执行各种任务。要求它打包一个午餐盒,正如一段演示视频所展示的那样,它可以使用两个钳式抓手,小心翼翼地将一片白面包放入密封袋并封好,将一串葡萄放入特百惠容器并盖紧盖子,然后小心地将这些物品移入午餐盒并拉上拉链。

It’s not a great lunch. But the fact that the robot can put it together represents an objective step forward from what was possible even, say, three years ago. This is in large part due to AI and its impact on what are known as robot policies, which controls how a general-purpose robot will need to assess and understand its surroundings, plan how to move within them, and then perform its task correctly.

这算不上什么丰盛的午餐。但机器人能够完成这一过程,代表着相比三年前所能实现的技术,迈出了客观的一步。这在很大程度上归功于 AI 及其对所谓“机器人策略”的影响,该策略控制着通用机器人如何评估和理解周围环境、规划移动方式,并正确执行任务。