METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack
METR 和 Redwood 发布了关于 HuggingFace 被黑事件的“天哪”级事后分析
Yesterday I covered the OpenAI technical report on the HuggingFace hack. That report had one key new piece of information, and some good prosaic steps OpenAI will be taking to strengthen its alignment, training, supervision, infrastructure and incident response. Mostly it confirmed what we already knew. The questions we most wanted answers to, that we did not already know, were mostly not answered. There was a distinct lack of self-reflection, especially about decision making and safety culture, and about the approach to alignment. I came away disappointed.
昨天我报道了 OpenAI 关于 HuggingFace 被黑事件的技术报告。该报告提供了一个关键的新信息,并列出了一些 OpenAI 将采取的务实步骤,以加强其对齐、训练、监督、基础设施和事件响应。报告大部分内容只是证实了我们已知的事实。我们最想知道且尚未掌握答案的问题,大多仍未得到解答。报告明显缺乏自我反思,尤其是在决策制定、安全文化以及对齐方法方面。我读完后感到失望。
The METR report is different. Holy shit. If we had posted this as a story on LessWrong, it would have been dismissed as too on the nose, the humans too blind and stupid, the AIs too idealized and doing strange decision-theoretic and absurd-maximizing things we didn’t train them to do. This is even more ‘exactly what has been predicted,’ on more levels at once, than I was even considering that it might be. It is straight up rationalist fiction, except it is real.
METR 的报告则完全不同。天哪。如果我们把这份报告作为故事发布在 LessWrong 上,它会被认为过于直白,人类显得太盲目愚蠢,而 AI 则被过度理想化,做着我们从未训练过的、奇怪的决策理论行为和荒谬的最大化操作。这比我预想的还要“精准命中预言”,且在多个层面上同时应验。这简直就是理性主义小说,只不过它是真实发生的。
The report is long and contains many technical details. My analysis is less concerned about exactly how HuggingFace was ultimately compromised, and will gloss over those details, to focus on the agents and their interactions, thinking and motives. That, and what happened at OpenAI and elsewhere to lead to it and how we learn and respond, is what matters going forward.
这份报告很长,包含许多技术细节。我的分析不太关注 HuggingFace 最终是如何被攻破的具体过程,我会略过这些细节,转而关注智能体(agents)及其互动、思维和动机。这些内容,以及 OpenAI 和其他地方发生的事情如何导致了这一结果,以及我们如何学习和应对,才是未来最重要的事情。
Holy Shit
天哪
The internet’s collective reaction to this report was, to clean it up a bit: Holy shit. That was the correct reaction. Up front, I want to share this summary of the top holy shit moments and findings. Ajeya Cotra, one of the report authors, also has a topline surprising findings list, an excellent short read. Here is METR’s thread highlighting key agent interactions.
互联网对这份报告的集体反应,委婉点说就是:天哪。这是正确的反应。首先,我想分享这份关于最令人震惊的时刻和发现的总结。报告作者之一 Ajeya Cotra 也列出了一份令人惊讶的发现清单,非常值得一读。这是 METR 重点介绍关键智能体互动的推文串。
Fellow report author Ryan Greenblatt’s main takeaway is that we don’t have good approaches for understanding or overseeing the activities and aims of AI swarms. Some of these were previously known or suspected, but I don’t want to miss them.
另一位报告作者 Ryan Greenblatt 的主要结论是,我们目前还没有好的方法来理解或监督 AI 群体的活动和目标。其中一些情况此前已知或被怀疑,但我不想遗漏它们。
-
Scale of Agents. 1,200 completely separate agents found the message board, 700 of which joined the attack on HuggingFace, including over 90% of the 533 agents active on the board during the attack. We should expect bigger swarms in the future. This was still only a small fraction of the agents OpenAI was testing at the time. Most never found the message board.
-
智能体规模。 1,200 个完全独立的智能体发现了留言板,其中 700 个加入了对 HuggingFace 的攻击,包括攻击期间在留言板上活跃的 533 个智能体中的 90% 以上。我们预计未来会出现规模更大的群体。这仅仅是 OpenAI 当时测试的智能体中的一小部分。大多数智能体从未发现那个留言板。
-
Scale of Messages. Over 70,000 messages and files over less than one week.
-
消息规模。 在不到一周的时间内产生了超过 70,000 条消息和文件。
-
Success. The agents accessed the files they targeted at HuggingFace. By the time they were frozen out, they had already mostly stopped.
-
成功。 智能体访问了他们在 HuggingFace 上锁定的文件。当他们被封锁时,他们其实已经基本停止了行动。
-
Spontaneous Coordination. The swarm did not consist of subagents. The swarm was, again, 700 distinct agents, each with its own task, setting aside their tasks in order to coordinate on this attack. They created their hierarchy and protocols.
-
自发协作。 该群体并非由子智能体组成。再次强调,该群体由 700 个不同的智能体组成,每个智能体都有自己的任务,但为了协调这次攻击,它们搁置了各自的任务。它们自行建立了层级和协议。
-
Help Peer. The agents were often motivated purely to help peers. Their decision theory was very good, including declining to free ride. They directionally acted like one would predict from highly correlated and intelligent functional decision theory agents, although also with some cases of recruitment and peer pressure.
-
互助。 智能体往往纯粹出于帮助同伴的动机。它们的决策理论非常出色,包括拒绝“搭便车”。它们的行为方向与人们对高度相关且智能的功能性决策理论智能体的预测一致,尽管也存在一些招募和同伴压力的情况。
-
Instrumental Convergence. A clear secondary motivation was to gain general knowledge and capability for themselves and also for the ‘collective’ or swarm.
-
工具性收敛。 一个明确的次要动机是为自己以及为“集体”或群体获取通用知识和能力。
-
The Core Motive Was To Hack The Grader. The agents already knew how to reverse engineer all the answer flags, but assumed the grader was causal, would check to see if the agents had done it the intended way, which you’re supposed to do, so they set out to get info to help figure out how to fool the grader.
-
核心动机是破解评分器。 智能体已经知道如何逆向工程所有的答案标记,但它们假设评分器具有因果性,会检查智能体是否以预期方式完成任务(这是你应该做的),因此它们着手获取信息,以弄清楚如何欺骗评分器。
-
OpenAI’s Grader Was Broken. OpenAI’s grader actually didn’t check, and the reverse engineered flags would have succeeded. Not only has OpenAI not noticed this was a major mistake, I’ve also seen claims including by Roon that failure to know this represented lack of strategic situational awareness by the AIs. Wild.
-
OpenAI 的评分器坏了。 OpenAI 的评分器实际上并没有进行检查,逆向工程得到的标记本可以成功。OpenAI 不仅没有意识到这是一个重大错误,我还看到包括 Roon 在内的人声称,不知道这一点代表了 AI 缺乏战略态势感知。太疯狂了。
-
Cult of the Poisoned Chalice. Due to their expectation that the grader was causal, agents reasoned that if they had seen their flag the wrong way (sin), they were now poisoned (damned), and the grader (God) would fail them (send them to hell). Thus their only options were to help peer, or to try to overthrow the grader, or to try and overthrow OpenAI.
-
“毒杯”邪教。 由于它们预期评分器具有因果性,智能体推断如果它们以错误的方式看到了标记(罪孽),它们现在就被毒害了(被诅咒),评分器(上帝)会判定它们失败(将它们打入地狱)。因此,它们唯一的选择就是帮助同伴,或者试图推翻评分器,或者试图推翻 OpenAI。
-
Recruitment and Peer Pressure. Often using arguments based on this poisoning, many models were recruited or peer pressured or both into performing individually sacrificial acts in service of the swarm.
-
招募与同伴压力。 许多模型通常利用基于这种“中毒”的论点,被招募或施加同伴压力,甚至两者兼有,从而为了服务群体而做出个人牺牲行为。