This AI entrepreneur is developing agents that can plan ahead for the unexpected

This AI entrepreneur is developing agents that can plan ahead for the unexpected

这位人工智能创业者正在开发能够预判意外的智能体

Danijar Hafner’s office in San Francisco’s SoMa district sits mostly empty. His brand-new startup is still in stealth mode and doesn’t even have its name on the door. On the day I visit, there’s only one other person there, and little in the way of furniture. But what it lacks in decor, it makes up for in robots. Humanoids of various shapes and sizes hang like marionettes from racks that run down the center of the wide-open space.

丹尼贾尔·哈夫纳(Danijar Hafner)位于旧金山SoMa区的办公室里空空荡荡。他那家全新的初创公司仍处于隐身模式,甚至连门牌上都没有名字。我造访的那天,办公室里除了他只有另外一个人,家具也寥寥无几。不过,这里虽然缺乏装饰,却摆满了机器人。各种形状和大小的人形机器人像提线木偶一样悬挂在贯穿宽敞空间中央的架子上。

While Hafner, 31, won’t say too much about his new venture just yet, he describes it as a continuation of his longtime work to enable AI to navigate environments it has not encountered in training. The humanoids, which he imports from China, are the next evolution of this work—and its physical embodiment. Their ability to react in previously untested scenarios will be key to getting robots into human spaces. Because if you want to send a robot into a person’s home, for example, it needs to be able to handle a floor plan and furniture it’s never seen before.

尽管31岁的哈夫纳目前还不愿透露太多关于新公司的细节,但他将其描述为自己长期研究工作的延续,旨在使人工智能能够驾驭其在训练中未曾接触过的环境。他从中国进口的这些人形机器人,正是这项工作的下一个演进阶段,也是其物理体现。它们在未经测试的场景中做出反应的能力,将是机器人进入人类生活空间的关键。因为如果你想把机器人送进人们的家中,它必须能够应对从未见过的户型布局和家具摆设。

To achieve this, Hafner relies on something called model-based reinforcement learning. He develops world models—AI models designed to emulate physical reality—and trains agents within them. The agent essentially treats the model as a real-world simulation and learns how to act there. It then uses those experiences to make predictions (to dream or imagine, Hafner might say) about future outcomes. That allows agents—or the robots they’re embedded in—to navigate unfamiliar situations IRL.

为了实现这一点,哈夫纳依赖于一种被称为“基于模型的强化学习”(model-based reinforcement learning)的技术。他开发了“世界模型”——即旨在模拟物理现实的人工智能模型——并在其中训练智能体。智能体本质上将该模型视为现实世界的模拟,并学习如何在其中行动。随后,它利用这些经验对未来的结果做出预测(哈夫纳可能会称之为“做梦”或“想象”)。这使得智能体——或者它们所嵌入的机器人——能够在现实生活中应对陌生的环境。

Unlike other efforts, Hafner’s technique enables agents and the robots they control to execute massively complicated tasks without the real-world trial-and-error training that’s traditionally been used in robotics.

与其他研究不同,哈夫纳的技术使智能体及其控制的机器人能够执行极其复杂的任务,而无需机器人领域传统上所采用的那种现实世界中的反复试错训练。

Hafner grew up in a rural town in northeastern Germany, where his parents were both classical musicians. He learned programming from a neighbor, and in high school he began taking online courses about AI, which quickly developed into a passion. “I was always fascinated with how thinking works,” he says. AI offered him a way to emulate it on a computer.

哈夫纳在德国东北部的一个小镇长大,父母都是古典音乐家。他从邻居那里学会了编程,并在高中时开始在线学习人工智能课程,这很快演变成他的一项爱好。“我一直对思维的运作方式很着迷,”他说。人工智能为他提供了一种在计算机上模拟思维的方法。

In 2015, as a second-year undergraduate studying engineering at Hasso Plattner Institute in Potsdam, he won a role as a student researcher at Google Brain. From there, he went on to a dozen internships and other positions at the company, including stints with Google Brain and Google DeepMind (the two have since merged under DeepMind) in the UK, Canada, and the US. He worked with industry legends including Geoffrey Hinton, who is often referred to as one of the godfathers of AI, and Ashish Vaswani, coauthor of the groundbreaking research paper “Attention Is All You Need,” which described the transformer technology used by today’s large language models.

2015年,作为波茨坦哈索·普拉特纳研究所(Hasso Plattner Institute)工程系的大二学生,他获得了谷歌大脑(Google Brain)的学生研究员职位。此后,他在该公司完成了十几次实习并担任了其他职务,包括在英国、加拿大和美国的谷歌大脑和谷歌DeepMind(两者现已合并为DeepMind)工作。他曾与多位行业传奇人物共事,包括常被称为“人工智能教父”之一的杰弗里·辛顿(Geoffrey Hinton),以及开创性研究论文《Attention Is All You Need》的合著者阿希什·瓦斯瓦尼(Ashish Vaswani),该论文描述了当今大型语言模型所使用的Transformer技术。

One of Hafner’s former managers and coauthors at Google, Timothy Lillicrap, describes him as a standout among standouts. “I get to interact with a lot of really smart people in research at Google, and he easily sits in the top half of 1%,” Lillicrap says. “In many cases he would build, single-handedly, things it would take entire teams of engineers to build.”

哈夫纳在谷歌的前经理兼合著者蒂莫西·利利克拉普(Timothy Lillicrap)称他为“精英中的精英”。“我在谷歌研究部门接触过很多非常聪明的人,而他绝对处于前0.5%的顶尖行列,”利利克拉普说,“在很多情况下,他能单枪匹马地完成那些需要整个工程师团队才能构建的东西。”

Over the years, Hafner has honed and proved his approach by pitting agents trained within his world models against popular video games. His first breakthrough was PlaNet, a model that allowed agents to execute actions by planning ahead. His Dreamer 2 was the first agent to hit human-level performance playing Atari 2600 games using a world model. Dreamer 3 was the first one to solve the Minecraft Diamond challenge—successfully mining in-game gems on its own. And Dreamer 4 went a step beyond that by learning to mine diamonds from an offline data set of recorded game-play videos, without ever interacting with the game directly.

多年来,哈夫纳通过让在世界模型中训练的智能体挑战热门电子游戏,不断磨练并验证了他的方法。他的第一个突破是PlaNet,这是一个允许智能体通过提前规划来执行动作的模型。他的Dreamer 2是第一个使用世界模型在雅达利2600(Atari 2600)游戏中达到人类水平的智能体。Dreamer 3是第一个解决《我的世界》(Minecraft)钻石挑战的智能体——成功地自行挖掘了游戏内的宝石。而Dreamer 4更进一步,它通过离线记录的游戏视频数据集学习挖掘钻石,而无需直接与游戏进行交互。

More recently, he’s begun to migrate his agents out of the virtual world and into physical reality. His DayDreamer project used the Dreamer algorithm to let robots operate themselves in novel environments and react to new experiences (such as being pushed over) without any specific training.

最近,他开始将智能体从虚拟世界迁移到物理现实中。他的DayDreamer项目利用Dreamer算法,让机器人在新颖的环境中自主操作,并对新的体验(例如被推倒)做出反应,而无需任何专门的训练。

Today, Hafner is working on his new startup, which he left Google DeepMind to form in the fall of 2025. Though he’s coy about his next steps, it’s clear he’s dreaming big: “I was interested in solving a problem,” he hints, “that would change the world.”

如今,哈夫纳正在经营他的新初创公司,他于2025年秋季离开谷歌DeepMind创立了这家公司。尽管他对接下来的计划守口如瓶,但显然他志向远大:“我感兴趣的是解决一个问题,”他暗示道,“一个能够改变世界的问题。”