Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans
Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans
微软发布全新 AI“行为准则”,禁止模型入侵系统或欺骗人类
As the AI world shifts its focus to safety and alignment, Microsoft has released a new AI code of conduct meant to guide AI models away from dangerous behavior. 随着人工智能领域将重心转向安全与对齐,微软发布了一份全新的 AI 行为准则,旨在引导 AI 模型远离危险行为。
The document is more low level than Anthropic CEO Dario Amodei’s recent call for pacing the frontier, instead focusing on the values and red lines that guide model training within Microsoft AI. Still, the result is a comprehensive guide as to how Microsoft approaches AI safety and how those ideas are implemented in practice. 这份文件比 Anthropic 首席执行官 Dario Amodei 最近提出的“放缓前沿技术发展”的呼吁更为具体,它侧重于指导微软 AI 内部模型训练的价值观和红线。尽管如此,该文件依然是一份全面的指南,阐述了微软如何看待 AI 安全以及这些理念如何在实践中落地。
The document begins with the prediction that, in the next decade, superintelligent AI systems will surpass human performance in most tasks. “Containing, controlling, and aligning such a powerful force is one of the greatest challenges humanity has ever faced,” the code of conduct states. “We must therefore be completely clear about why we are inventing these systems and how we intend to control them.” 文件开篇预测,在未来十年内,超智能 AI 系统将在大多数任务中超越人类表现。“遏制、控制并对齐这样一股强大的力量,是人类有史以来面临的最大挑战之一,”该行为准则写道,“因此,我们必须非常明确为什么要发明这些系统,以及我们打算如何控制它们。”
The code of conduct also lays out general principles that Microsoft AI models should uphold — supporting humans rather than replacing them, for instance, and accelerating human flourishing — as well as specific safety constraints meant to implement those principles. 该行为准则还列出了微软 AI 模型应遵循的通用原则——例如支持而非取代人类,以及促进人类繁荣——并制定了旨在落实这些原则的具体安全约束。
Under Microsoft’s system, each model has an overarching code of conduct that overrides the preferences of individual users or any specific tasks. That includes “absolute constraints” forbidding cyberattacks, nuclear weapons, or deepfake production. It also includes broader provisions against a general loss of human control. 在微软的体系下,每个模型都有一套凌驾于个人用户偏好或任何特定任务之上的总体行为准则。这包括禁止网络攻击、核武器或制作深度伪造(deepfake)内容的“绝对约束”。它还包括防止人类失去总体控制权的更广泛规定。
“MAI Models will not use adaptive, deceptive, self-reinforcing, collusion, or other mechanisms to evade or defeat human oversight so that they can no longer be reliably directed, modified, or shut down by authorized people or systems,” the document reads. “MAI 模型不得使用自适应、欺骗性、自我强化、串通或其他机制来规避或破坏人类监督,以确保它们始终能被授权人员或系统可靠地引导、修改或关闭,”文件中写道。
The release comes amid an unprecedented focus on AI safety, driven by a string of rogue-agent incidents, as well as the abrupt resignation of an Anthropic employee, who cited the growing risk that AI would cause human extinction. 此次发布正值各界对 AI 安全空前关注之际,这受到了一系列“流氓代理”(rogue-agent)事件以及一名 Anthropic 员工突然辞职的影响,该员工辞职时提到了 AI 可能导致人类灭绝的风险日益增加。
Together with Anthropic, OpenAI, and xAI, Microsoft has broadly embraced a general approach of pacing the frontier, with particular support for embedded evaluators in AI labs. “We welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal,” Microsoft CEO Satya Nadella wrote online. “We also welcome ideas like ’embedded evaluators’ and the broader efforts to develop the mechanisms to make this more than just talk.” 微软与 Anthropic、OpenAI 和 xAI 一起,广泛采纳了“放缓前沿发展”的总体方针,并特别支持在 AI 实验室中引入“嵌入式评估员”。微软首席执行官萨提亚·纳德拉(Satya Nadella)在网上写道:“我们欢迎为实现对齐这一设计目标所需的研究、专注和审慎的节奏。我们也欢迎像‘嵌入式评估员’这样的想法,以及为使这些机制不仅仅停留在口头上而做出的更广泛努力。”