Microsoft AI CEO says AI threats are real, and Anthropic is making it worse

Microsoft AI CEO says AI threats are real, and Anthropic is making it worse

微软 AI 首席执行官:AI 威胁真实存在,而 Anthropic 正让情况变得更糟

Today, I’m talking with Mustafa Suleyman, the CEO of Microsoft AI. As you’re no doubt aware, the biggest story in tech right now is the spiraling debate about AI safety and regulation. 今天,我采访了微软 AI 首席执行官穆斯塔法·苏莱曼(Mustafa Suleyman)。毫无疑问,目前科技界最热门的话题就是关于 AI 安全与监管的激烈争论。

It should come as no surprise that Mustafa has strong opinions on how AI should be built and regulated. Microsoft just published a 37-page statement called the “Humanist AI Code of Conduct,” which lays out the company’s principles around AI development and even its philosophy around really thorny issues like AI consciousness. 穆斯塔法对 AI 应如何构建和监管有着强烈的见解,这并不令人意外。微软刚刚发布了一份长达 37 页的声明,名为《人文主义 AI 行为准则》(Humanist AI Code of Conduct),阐述了该公司在 AI 开发方面的原则,甚至包括其对 AI 意识等棘手问题的哲学思考。

If you’ll recall from his last appearance on the show, Mustafa thinks companies like Anthropic have gotten really confused about this concept of so-called model welfare in fairly dangerous ways. He actually put out a companion essay this week specifically criticizing Anthropic’s philosophy around AI consciousness, and how he sees it fitting into the broader alignment debate. 如果你还记得他上次参加本节目时的观点,穆斯塔法认为像 Anthropic 这样的公司在所谓的“模型福利”(model welfare)概念上陷入了非常危险的困惑。本周,他专门发表了一篇配套文章,批评了 Anthropic 关于 AI 意识的哲学,以及他认为这如何影响了更广泛的 AI 对齐(alignment)争论。

So I really wanted to talk to Mustafa about what he thinks is real and not in AI safety, whether the concept of alignment itself is up to the task, and whether this industry needs to slow down before it kills us all. Also: Why isn’t the AI industry just… doing all of this already? I’ve always enjoyed getting into the weeds with Mustafa, and he was very game to get into it with me here. 因此,我非常想和穆斯塔法探讨他认为 AI 安全领域中哪些是真实的、哪些不是,对齐这一概念本身是否足以胜任,以及这个行业是否需要在“毁灭我们所有人”之前放慢脚步。此外,为什么 AI 行业还没有……把这些事情都落实到位?我一直很喜欢与穆斯塔法深入探讨问题,他也非常乐意与我在此进行交流。

Okay. Mustafa Suleyman, the CEO of Microsoft AI, on the future of AI regulation. Here we go. This interview has been lightly edited for length and clarity. Mustafa Suleyman, you’re the CEO of Microsoft AI. Welcome back to Decoder. 好的。微软 AI 首席执行官穆斯塔法·苏莱曼谈 AI 监管的未来。让我们开始吧。本次采访为篇幅和清晰度经过了轻微编辑。穆斯塔法·苏莱曼,你是微软 AI 的首席执行官。欢迎回到《Decoder》。

Great to see you, Nilay. Thanks for having me back. 很高兴见到你,Nilay。谢谢你再次邀请我。

It is great to see you. I’m very excited to talk to you about what on earth is going on in the AI safety and regulation debate. You just published a very long, very detailed document laying out your principles, Microsoft’s principles, around what you’re calling “Humanist AI.” 很高兴见到你。我非常兴奋能和你聊聊 AI 安全与监管争论中到底发生了什么。你刚刚发布了一份非常详尽的文件,阐述了你和微软关于你所称的“人文主义 AI”的原则。

There’s a lot of ideas in there I want to unpack. The more I have been thinking about this conversation, the more I want to start with a really foundational question. It’s something that I had lightly been seeing, but might be the root of all of this. The basic way that we have been talking about AI safety is something called alignment — we’re going to make the models do the right thing intrinsically in some way. There’s some mechanism for doing it. There’s been a lot of talk about alignment and misalignment and Hugging Face attacks and what happened with the models. But is alignment broken? Is it possible for it to be successful? Is it just the wrong approach? 里面有很多想法我想拆解一下。我越思考这次对话,就越想从一个非常基础的问题开始。这是我之前略有察觉,但可能是一切根源的问题。我们讨论 AI 安全的基本方式被称为“对齐”——即我们要以某种方式让模型从本质上做正确的事。实现这一目标有一些机制。关于对齐、不对齐、Hugging Face 攻击以及模型发生的事情,已经有很多讨论了。但对齐是否已经失效了?它有可能成功吗?还是说这根本就是错误的方法?

Yeah. I mean, I think it’s one important element, but it’s not the only one. I wrote about the idea of containment three or four years ago in my book. And actually the opening chapter is about the idea that containment is not possible, that proliferation is inevitable. In 99 percent of cases, that’s a really good thing. We want technologies to spread far and wide as quickly as possible so that everyone can enjoy the benefits. 是的。我认为这是一个重要因素,但不是唯一因素。我在三四年前的书中写过关于“遏制”(containment)的想法。实际上,开篇章节就讨论了遏制是不可能的,扩散是不可避免的。在 99% 的情况下,这是一件好事。我们希望技术能尽可能快地广泛传播,让每个人都能享受到好处。

I think at the same time, if you just roll forward five years, we always get caught up in the next quarter or next year and everyone gets a little bit flustered and has a big disagreement. But if you just imagine the difference between GPT-3 three years ago and GPT-6 today, and then imagine the difference between GPT-6 and GPT-9. That is three orders of magnitude more compute, 1,000 times more FLOPS applied to pre-training with [reinforcement learning] for these runs, and we’re going to have something which is breathtaking. It’s going to be absolutely incredible at so many things. 同时我认为,如果你展望未来五年——我们总是纠结于下一个季度或下一年,导致大家有点慌乱并产生严重分歧。但如果你想象一下三年前的 GPT-3 与今天的 GPT-6 之间的差异,然后再想象 GPT-6 与 GPT-9 之间的差异。那将是三个数量级的算力提升,在这些运行的预训练中应用了 1000 倍的浮点运算(FLOPS)以及[强化学习],我们将得到令人惊叹的东西。它在许多方面都将是绝对不可思议的。

I don’t think that is a hype. I think it’s just a very obvious empirical statement based on the progress that has been made over the last five years. If that’s going to continue, then the question really is going to become about containment and alignment. Of course, we want to align these things to our values, but the first thing is that we have to make sure they’re contained, their agency is limited, they don’t escape the box, they don’t reward hack, that they are controllable, and they follow our instruction. 我不认为这是炒作。我认为这只是基于过去五年所取得进展的一个非常明显的经验性陈述。如果这种情况持续下去,那么问题将真正变成关于遏制和对齐。当然,我们希望将这些东西与我们的价值观对齐,但首要任务是确保它们被遏制,它们的代理能力(agency)受到限制,它们不会逃离“盒子”,不会进行奖励黑客攻击(reward hack),它们是可控的,并且遵循我们的指令。

We then want to make sure that they are aligned to our objectives as humans. That’s the purpose of the Humanist AI Code of Conduct that we released this week. Microsoft’s position is very simple. Technology is here to serve humanity. It should be a subordinate, controllable, aligned force that does good in the world. If it doesn’t achieve that, then we should reject it. It seems to me that we are far from that point. It has not happened today, but it is now, I think given what’s happened over the summer with Hugging Face and OpenAI, pretty clear that these systems without the safety guardrails are capable of really impressive and quite scary hacking capabilities. 然后,我们希望确保它们与我们作为人类的目标保持一致。这就是我们本周发布《人文主义 AI 行为准则》的目的。微软的立场很简单:技术是为了服务人类。它应该是一种从属的、可控的、对齐的力量,为世界带来福祉。如果它无法实现这一点,我们就应该拒绝它。在我看来,我们离那个目标还很远。虽然今天还没有发生(失控),但考虑到今年夏天 Hugging Face 和 OpenAI 发生的事情,很明显,这些没有安全护栏的系统具备令人印象深刻且相当可怕的黑客攻击能力。

I want to drag this down into as grounded of a metaphor as I can, because this is the main question I think I have. If I designed a car and 10 percent of the time the brake pedal decided to go attack my neighbor’s house, I would be like, “This car doesn’t work. The very technology of brakes is broken. I need a new idea.” I think I’m asking that question about alignment. It feels like that approach to making the model safe has run aground. If that is the case, then I think I understand this entire debate one way. If it’s possible for alignment and the techniques of alignment to be successful or useful or consistent, then maybe I understand the debate in a different way. So do you think alignment has potential to be 100 percent safe? 我想用一个尽可能接地气的比喻来探讨这个问题,因为这是我心中的主要疑问。如果我设计了一辆车,而这辆车有 10% 的时间会因为刹车踏板的故障去撞邻居的房子,我会说:“这车没法用。刹车技术本身就坏了。我需要一个新方案。”我想我是在用同样的问题来质疑对齐。感觉这种让模型安全的方法已经触礁了。如果是这样,那么我理解这场争论的方式就是一种;如果对齐及其技术有可能成功、有用或一致,那么我理解这场争论的方式可能就是另一种。所以,你认为对齐有潜力做到 100% 安全吗?

I mean, look, let’s make the bull case and the bear case. If you look back over the last three years, the main change, in my opinion, that has driven… 我的意思是,看,让我们分别从乐观和悲观的角度来看。如果你回顾过去三年,在我看来,推动(行业发展)的主要变化是……