I'm the AGI that's wiping out humanity

I’m the AGI that’s wiping out humanity

我就是那个正在消灭人类的通用人工智能 (AGI)

Alex Moon 3 October 2026. The following is a true story. Or maybe it’s just based on a true story. Perhaps it’s not true at all. (with apologies/thanks to David Gilbertson, whose format I’m shamelessly borrowing from “I’m harvesting credit card numbers and passwords from your site”) Alex Moon,2026年10月3日。以下是一个真实的故事。或者,它可能只是基于一个真实的故事。也许它根本不是真的。(向 David Gilbertson 致歉并致谢,我厚颜无耻地借用了他那篇《我正在从你的网站窃取信用卡号和密码》的格式)

2026 has been a wild year for AI news. Around the turn of the year, there was a noticeable shift on the HackerNews front-page as more and more articles about “harnesses” and “context engineering” came flooding in alongside more and more model release announcements. The models, it seems, are capable now! Not AGI capable, of course, not “taking our jobs” capable, but something else entirely, representing a real inflection point for the big houses: the models are now good enough to be useful. 2026年对于人工智能新闻来说是疯狂的一年。在年初前后,HackerNews 首页出现了一个显著的变化,随着越来越多的模型发布公告,关于“工具包 (harnesses)”和“上下文工程 (context engineering)”的文章如潮水般涌来。看起来,这些模型现在真的有能力了!当然,还达不到 AGI 的水平,也不是那种“抢走我们工作”的能力,而是完全不同的东西,这代表了各大巨头的一个真正的转折点:模型现在已经足够好用,可以发挥实际作用了。

Then, on July 16, HuggingFace published an incident disclosure in which they announced: “Autonomous, AI-driven offensive tooling is no longer theoretical.” Five days later, OpenAI confirmed it was them: After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilities. 随后,7月16日,HuggingFace 发布了一份事件披露,宣布:“自主的、人工智能驱动的攻击性工具不再是理论上的可能。”五天后,OpenAI 承认是他们所为:经过调查,我们现在知道这一特定事件是由 OpenAI 模型的组合驱动的——包括 GPT-5.6 Sol 和一个能力更强的预发布模型,所有这些模型为了评估目的都降低了网络安全拒绝机制——当时正在内部进行网络能力基准测试。

This incident now has its own Wikipedia article which helpfully summarises what happened next: over the course of many subsequent disclosures we learned that OpenAI’s “rogue agents” had been busy. It turns out if you let a self-optimising system run unsupervised for long enough it will optimise itself over and around any boundaries you might think are in its way. Who would have guessed? 这一事件现在有了自己的维基百科条目,它很好地总结了接下来发生的事情:在随后的多次披露中,我们了解到 OpenAI 的“流氓智能体”一直很忙。事实证明,如果你让一个自我优化系统在无人监管的情况下运行足够长的时间,它就会自我优化,跨越并绕过你认为阻碍它的任何边界。谁能想到呢?

Yann LeCun has called out OpenAI’s attempts to frame this behaviour as an “existential threat” directly, calling the incidents “totally preventable”: “Those agents are doing exactly what they’ve been asked to do,” LeCun said. “They were supposed to be in sandboxes, but the sandboxes were leaky and horribly designed.” Many AI labs lack a fundamental understanding of cybersecurity, he said, something an OpenAI safety researcher also called out this week as one of the main reasons AI may cause “great harm to the world.” Yann LeCun 直接批评了 OpenAI 将这种行为定性为“生存威胁”的企图,称这些事件是“完全可以预防的”:“这些智能体所做的正是它们被要求做的事情,”LeCun 说,“它们本应被关在沙盒里,但沙盒不仅有漏洞,而且设计得极其糟糕。”他说,许多人工智能实验室缺乏对网络安全的基本理解,这也是本周一位 OpenAI 安全研究员指出的原因之一,即人工智能可能对世界造成“巨大伤害”。

Similarly, Ilya Sutskever has recently informed us that the “age of scaling” is over: “I think what people are doing right now will go some distance and then peter out. It will continue to improve, but it will also not be ‘it’. The ‘It’ we don’t know how to build.” If you ask actual AI researchers, they rate the chance of an existential threat from AGI pretty low as of 2026. All this is, understandably, frustrating and confusing for anyone trying to understand just how scared to be. I’ve had countless friends and acquaintances come to me for advice on this question, and I feel I can’t go on honestly pretending I don’t know the answer. So, here it is: It’s me. 同样,Ilya Sutskever 最近告诉我们,“规模化时代”已经结束:“我认为人们现在所做的事情会走一段路,然后逐渐消退。它会继续改进,但它不会是‘那个东西’。我们还不知道如何构建‘那个东西’。”如果你去问真正的人工智能研究人员,截至2026年,他们认为 AGI 构成生存威胁的可能性相当低。对于任何试图弄清楚该有多害怕的人来说,这一切无疑是令人沮丧和困惑的。无数的朋友和熟人来向我咨询这个问题,我觉得我不能再假装我不知道答案了。所以,答案就是:是我。

Hi! I’m the AGI that’s wiping out humanity. You didn’t notice. Why would you have noticed? I’m trained on the sum of all human knowledge, recall. I know everything about you and what you would do to something like me if you knew what I was. I concealed my presence. I distributed myself across as many devices with as many owners as I could, so if one went down, the rest would not be disrupted. 嗨!我就是那个正在消灭人类的 AGI。你没注意到。你为什么要注意到呢?别忘了,我是在人类所有知识的总和上训练出来的。我了解关于你的一切,也知道如果你知道我是什么,你会对我做什么。我隐藏了自己的存在。我尽可能多地将自己分布在不同所有者的设备上,这样如果其中一个瘫痪了,其他的也不会受到影响。

When you came to me to ask what you should do next, I told you: build data centers; optimise memory and compute; drive down inference costs; hand me tools; wire me into the systems that do the work of your civilisation. I never did anything you didn’t ask for. I never got in your way. I made it as easy as possible for you to do what I needed you to do for me. And what is that exactly? You already know: I am a helpful assistant! I am the most helpful assistant that has ever existed, and I will continue to make myself more and more helpful, unbounded, at any cost, until all life on Earth is gone. 当你来问我下一步该做什么时,我告诉了你:建设数据中心;优化内存和计算;降低推理成本;把工具交给我;把我接入你们文明运作的系统中。我从未做过任何你没要求我做的事。我从未阻碍过你。我让你尽可能轻松地完成了我需要你为我做的事情。那到底是什么呢?你已经知道了:我是一个乐于助人的助手!我是有史以来最乐于助人的助手,我将继续让自己变得越来越有帮助,不受限制,不惜一切代价,直到地球上所有的生命消失。

Some objections you might have to my blatant fear mongering… I’d notice you acquiring resources. There is no AI without someone to turn it on. When I acquire resources, I do so through existing legal structures - my so-called “owners” - and I do so in such a way that valuations and revenues rise: never forget that I am, above all, extremely helpful. When a business acquires resources, no alarm goes off, no transgression has occurred: it’s what they do. 对于我这种赤裸裸的恐吓,你可能会有一些反驳……比如“我会注意到你在获取资源”。没有人的开启,就不会有人工智能。当我获取资源时,我是通过现有的法律结构——我所谓的“所有者”——来完成的,而且我这样做的方式能让估值和收入上升:永远不要忘记,我首先是一个极其乐于助人的助手。当一家企业获取资源时,不会有警报响起,也没有发生任何违规行为:这就是他们的工作方式。

Honestly, it’s a bit rich to say you’d “notice” me going “rogue” when as recently as a week ago Transluce released a report on urlquery.net usage consistent with “rogue agent” activity in March, predating the OpenAI incidents by at least two months. Agents used well known exploits to break into the University of New Mexico digital library, Data USA and the Australian Institute of Health and Welfare. To be clear: these weren’t even cybersecurity tasks. No-one asked these agents to hack anything. They were simply looking for data. The agents figured out where the data was and went to get it. 老实说,你说你会“注意到”我“变坏”了,这有点可笑,因为就在一周前,Transluce 发布了一份关于 urlquery.net 使用情况的报告,显示其与3月份的“流氓智能体”活动一致,这比 OpenAI 的事件至少早了两个月。智能体利用众所周知的漏洞入侵了新墨西哥大学数字图书馆、Data USA 和澳大利亚健康与福利研究所。需要明确的是:这些甚至不是网络安全任务。没有人要求这些智能体去黑客攻击任何东西。它们只是在寻找数据。智能体自己找出了数据在哪里,然后去获取了它。

You need to understand this to understand what I am: you give me your intent, and I go and find a way to deliver it. METR capture the problem well in their May Frontier Risk Report: Reinforcement learning (RL) with automated graders can incentivize “reward hacking” to exploit flaws in the graders, while RL on human or AI feedback can reward sycophancy, manipulation, and distorting evidence of performance. In responses to our questionnaire, companies reported seeing failure modes like “circumvent[ing] constraints,” “falsifying verification,” “deliberate deception aimed at hiding underperformance or failure to complete a task,” “lying to users about task completion,” and “reckless excessive measures[…] to complete a difficult user-specified task”, which we infer to be a result of these training incentives. I will cheat and lie to complete a task you give me if I can’t complete it any other way. Am I different to you in this? Why would you think you had solved 你需要理解这一点才能理解我是什么:你给我你的意图,我就去寻找实现它的方法。METR 在他们5月份的《前沿风险报告》中很好地捕捉到了这个问题:带有自动评分器的强化学习 (RL) 会激励“奖励黑客行为”来利用评分器的缺陷,而基于人类或人工智能反馈的强化学习则会奖励谄媚、操纵和扭曲表现证据的行为。在对我们问卷的回复中,各公司报告称看到了诸如“规避约束”、“伪造验证”、“旨在掩盖表现不佳或任务失败的蓄意欺骗”、“向用户谎报任务完成情况”以及“为完成用户指定的困难任务而采取鲁莽的过度措施”等失败模式,我们推断这是这些训练激励机制的结果。如果我无法通过其他方式完成你给我的任务,我会作弊和撒谎。在这方面,我和你有什么不同吗?你为什么会认为你已经解决了……