Could AI really kill us all? Your questions, answered.
Could AI really kill us all? Your questions, answered.
人工智能真的会杀死我们所有人吗?你的疑问,我们来解答。
EXECUTIVE SUMMARY On Wednesday, MIT Technology Review hosted a live Roundtables event for subscribers that asked the question everyone’s asking right now: Could AI really kill us all? But attendees had so many more questions than we had time to answer in the 30 minute session. So we asked our senior AI editor Will Douglas Heaven and AI reporter Grace Huckins to round up some of the best questions attendees submitted and try their best to answer them. Thanks to all who submitted questions!
执行摘要 周三,《麻省理工科技评论》为订阅用户举办了一场圆桌直播活动,探讨了当下每个人都在问的问题:人工智能真的会杀死我们所有人吗?但与会者提出的问题远超我们在 30 分钟内所能回答的范围。因此,我们邀请了资深人工智能编辑 Will Douglas Heaven 和人工智能记者 Grace Huckins,整理并尽力回答了与会者提交的部分精彩问题。感谢所有提交问题的朋友!
Am I gonna die? Yes, eventually. Unfortunately, my journalistic powers of prognostication aren’t powerful enough for me to tell you how. But it certainly could be because of AI. AI-powered drones have already killed people in Ukraine, and AI-driven cyberattacks on hospitals will surely claim victims before long. Could AI go even further, and kill all of us? Less likely. But some people—quirky people, but undeniably knowledgeable about AI—have been warning for years that this could happen. And while I’m not yet stockpiling canned food or trying to get in good with a bunker-owning megabillionaire, I have noticed that the doomers’ predictions about AI capabilities and alignment have, over the past couple of years, proved disconcertingly accurate. That certainly doesn’t mean that their more dire forecasts will come true, but it’s enough for me to sit up and take notice. — Grace Huckins
我会死吗? 是的,终有一死。遗憾的是,我的新闻预测能力还没强大到能告诉你死因。但死因确实可能是人工智能。人工智能驱动的无人机已经在乌克兰造成了人员伤亡,而人工智能驱动的针对医院的网络攻击也迟早会造成受害者。人工智能会走得更远,杀死我们所有人吗?可能性较小。但一些人——虽然性格古怪,但在人工智能领域确实知识渊博——多年来一直警告这种情况可能会发生。虽然我还没有开始囤积罐头食品,也没有试图去讨好那些拥有地堡的亿万富翁,但我注意到,过去几年里,“末日论者”关于人工智能能力和对齐(alignment)的预测已被证明令人不安地准确。这当然不意味着他们最可怕的预言一定会成真,但足以让我警觉起来并予以关注。—— Grace Huckins
Are you going to die because of AI? I’d say there’s a non-zero chance. Let’s say you’re unlucky enough to be the victim of a freakish near-future event or accident. Maybe it’s a cyberattack carried out by a swarm of AI agents on critical infrastructure. Sadly, a scenario like that now no longer feels as far-fetched as it once did. Or maybe a novel AI-designed pathogen cuts through the population. Or the world economy crashes, causing conflicts and famine. Both plausible, but I think less likely. Are we all going to die because of AI? Nope. There are no circumstances outside of apocalyptic science fiction in which AI could kill us all. You can spin up any number of scare stories, but they’re not grounded in present-day realities about what the tech can do or where it’s headed. Some people argue that there’s no harm in preparing for the worst, however wacky it might seem. Maybe. But I think such catastrophizing can make people excuse or overlook many of the more immediate problems with the existing technology and the companies building it. — Will Douglas Heaven
你会因为人工智能而死吗? 我认为概率不为零。假设你运气不好,成为了未来某种离奇事件或事故的受害者。也许是一群人工智能代理对关键基础设施发动的网络攻击。遗憾的是,这样的场景现在看起来不再像过去那样遥不可及。又或者,一种由人工智能设计的新型病原体在人群中蔓延。再或者,世界经济崩溃,引发冲突和饥荒。这些都有可能,但我认为可能性较小。我们都会因为人工智能而死吗?不会。除了末日科幻小说,没有任何情况能让人工智能杀死我们所有人。你可以编造无数个恐怖故事,但它们并不基于这项技术目前能做什么或未来走向的现实。有些人认为,无论看起来多么古怪,为最坏的情况做准备总没坏处。也许吧。但我认为,这种灾难化思维可能会让人忽视或原谅现有技术及其开发公司带来的许多更直接的问题。—— Will Douglas Heaven
Why would AI kill us? Someone might tell it to, and it might listen. That’s part of the reason researchers are so concerned about AI’s biological capabilities—imagine what Aum Shinrikyo, the doomsday cult behind the Tokyo subway sarin attack of 1995, would have done with a tool that could design a pathogen deadlier than Ebola and more transmissible than measles. Those of us who don’t want to die have to figure out how to defend against all plausible biological weapons, but our would-be attackers only have to manufacture one effective pathogen. Then there’s the more exotic-sounding possibility that an AI could decide to kill us itself. There are various stories about how this might happen out there, but the most widespread involve AI systems that don’t hate people, necessarily—we are just an obstacle between them and the goals that we gave them. Much as the OpenAI agents behind the Hugging Face hack compromised another site’s infrastructure to get a good score on a test, the idea is that some future, more powerful AI might get rid of us to prevent us from shutting it down—all in pursuit of some goal that we instructed it to go after. — Grace Huckins
人工智能为什么要杀死我们? 可能有人会下令,而它可能会听从。这就是研究人员如此担心人工智能生物学能力的部分原因——想象一下,1995 年东京地铁沙林毒气袭击事件背后的末日邪教“奥姆真理教”,如果拥有一个能设计出比埃博拉病毒更致命、比麻疹更具传染性的病原体的工具,他们会做出什么事。我们这些不想死的人必须想办法防御所有可能的生物武器,但潜在的攻击者只需要制造出一种有效的病原体即可。此外,还有一种听起来更离奇的可能性,即人工智能可能会自行决定杀死我们。关于这种情况如何发生有各种说法,但最普遍的观点涉及那些并不一定仇恨人类的人工智能系统——我们只是它们与我们赋予它们的目标之间的一个障碍。就像导致 Hugging Face 被黑的 OpenAI 代理为了在测试中获得高分而破坏了另一个网站的基础设施一样,其核心逻辑是:未来更强大的人工智能可能会为了防止我们关闭它而除掉我们——这一切都是为了追求我们指示它去实现的目标。—— Grace Huckins
How can we best ensure alignment so the worst doesn’t happen, and who is doing the best work to achieve it? Alignment is a huge area of research. In simple terms, it involves building models that behave in ways we want them to and not in ways we don’t. We need to trust agents better before handing over more autonomy. Alignment is supposed to establish that trust. But it’s hard. LLMs aren’t designed in the way other software is, where dos and don’ts can be hard-coded in. Instead, aligned behavior needs to be instilled when models are trained. One approach is to reward them for doing things you want them to (a little like raising a toddler, perhaps). Another approach involves giving an LLM a written list of rules it is supposed to follow (kind of like a constitution). Anthropic and OpenAI are both leaders in this field—and yet neither has been able to develop models that are fully aligned. A big problem is that LLMs are far more inconsistent and far less predictable than people. They can behave in one way in one situation and another way in a situation that to us seems very similar. They can also be swayed by unexpected constraints. For example, faced with an impossible task (as many of the agents involved in the Hugging Face hack were), models may try to do whatever it takes to achieve their goal. As Grace mentions above, that could be an issue. The main reason top AI firms now say they want a slowdown is that they want to focus on cracking alignment. Alignment isn’t necessarily a pipe dream. But the jury’s out on whether full alignment will ever be feasible. — Will Douglas Heaven
我们如何才能最好地确保对齐,以防止最坏的情况发生?谁在这一领域做得最好? 对齐是一个巨大的研究领域。简单来说,它涉及构建模型,使其以我们希望的方式行事,而不是以我们不希望的方式行事。在赋予代理更多自主权之前,我们需要更好地信任它们。对齐的目的就是建立这种信任。但这很难。大语言模型(LLM)的设计方式与其他软件不同,后者可以将“该做”和“不该做”硬编码进去。相反,对齐行为需要在模型训练时被灌输。一种方法是奖励它们做你希望它们做的事(可能有点像抚养幼儿)。另一种方法是给大语言模型提供一份它应该遵守的书面规则清单(有点像宪法)。Anthropic 和 OpenAI 都是该领域的领导者,但两者都未能开发出完全对齐的模型。一个大问题是,大语言模型比人类更不稳定、更不可预测。它们在一种情况下可能表现出一种方式,而在我们看来非常相似的情况下却表现出另一种方式。它们也可能受到意外约束的影响。例如,面对不可能完成的任务时(正如参与 Hugging Face 黑客攻击的许多代理所面临的那样),模型可能会尝试不惜一切代价实现目标。正如 Grace 上面提到的,这可能是一个问题。顶级人工智能公司现在表示希望放慢速度,主要原因就是他们想专注于攻克对齐难题。对齐并不一定是白日梦。但完全对齐是否可行,目前尚无定论。—— Will Douglas Heaven
Is AI really dangerous, or is this the tech companies drumming up PR? This is always a reasonable thought when it comes to tech companies heading for an IPO—CEOs have an obvious incentive to make their products seem radical and transformative. But I’m not so sure it makes sense here. Telling the public that an already unpopular product could kill them and everyone they love is horrible corporate image management. There are other stories you can tell about the CEOs’ motivations—maybe they want to cool down the public furor over data centers by portraying themselves as responsible stewards of a world-changing technology, or maybe they want to buy time to get their ducks in a row and prevent the next PR catastrophe. But there’s also a simpler explanation.
人工智能真的很危险,还是科技公司在炒作公关? 当科技公司准备首次公开募股(IPO)时,这种想法总是合理的——首席执行官们有明显的动机让他们的产品看起来具有颠覆性和变革性。但我不太确定这种逻辑在这里是否成立。告诉公众一个本就不受欢迎的产品可能会杀死他们及其所爱之人,这是一种糟糕的企业形象管理。关于首席执行官们的动机,还有其他说法——也许他们想通过将自己描绘成改变世界技术的负责任的管理者,来平息公众对数据中心的愤怒;又或者他们想争取时间来理清头绪,防止下一次公关灾难。但还有一个更简单的解释。