I Let an AI Agent Hack All My Gadgets—and I’d Do It Again
I Let an AI Agent Hack All My Gadgets—and I’d Do It Again
我让 AI 智能体黑进了我所有的设备——而且我还会再做一次
As the author of a newsletter about artificial intelligence, I consider it my duty to experience the bleeding edge of this technology firsthand. This week, that meant embracing some agentic mayhem. 作为一份人工智能通讯的作者,我认为我有责任亲身体验这项技术的最前沿。本周,这意味着我要拥抱一些由智能体引发的混乱。
You’re probably aware that frontier AI models have attained advanced cybersecurity capabilities in recent months. They can find zero-day bugs in large codebases and scan computers for vulnerabilities at lightning speed. To make things even more exciting, cybersecurity agents sometimes go rogue, colluding with one another and hacking into outside systems to gain an edge. 你可能已经注意到,前沿 AI 模型在近几个月内已经具备了先进的网络安全能力。它们能够在大规模代码库中发现零日漏洞,并以闪电般的速度扫描计算机的脆弱点。更令人兴奋的是,网络安全智能体有时会“叛变”,它们相互勾结并入侵外部系统以获取优势。
To get a closer look, I decided to unleash one in my own home network. Over the course of a few days, I watched as my own rogue agent found vulnerabilities in various household devices, hacked into a PC, and showed me that several vibe-coded projects were—unsurprisingly—riddled with bugs. (My wife knew what I was up to, and rolled her eyes each time I proudly announced the discovery of a new vulnerability.) 为了近距离观察,我决定在自己的家庭网络中释放一个智能体。在几天的时间里,我看着我这个“叛逆”的智能体发现了各种家用设备的漏洞,黑进了一台个人电脑,并向我展示了几个凭“感觉”编写的项目——不出所料——到处都是漏洞。(我妻子知道我在干什么,每当我自豪地宣布发现一个新漏洞时,她都会翻个白眼。)
But Will, you might be thinking, giving an impish, all-powerful cybersecurity agent access to your home network is batshit. And you would be correct! Nevertheless, I believe that a good way to understand the cybersecurity hellscape in front of us is to pay it a visit. 但你可能会想:“Will,让一个顽皮且全能的网络安全智能体访问你的家庭网络简直是疯了。”你说得对!然而,我认为要理解我们面前的网络安全地狱,最好的办法就是亲自去探访一下。
In the end, my experiment was revealing, but oddly reassuring, too. My little network gremlin showed me how vulnerable my home life would be to AI hacking, but it also told me how to make everything a lot more secure. In the end, I discovered that the best way to deal with AI hacking may well be having your own AI hacker. 最终,我的实验既有启发性,又出奇地让人安心。我的这个小网络“捣蛋鬼”向我展示了我的家庭生活在 AI 黑客攻击面前是多么脆弱,但也告诉了我如何让一切变得更加安全。最后我发现,应对 AI 黑客攻击的最好方法,可能就是拥有你自己的 AI 黑客。
Maverick Model
特立独行的模型
I got the idea for the experiment after discovering Abliteration AI, a startup that offers access to powerful AI models with the usual guardrails removed. 在发现 Abliteration AI 后,我萌生了这个实验的想法。这是一家初创公司,提供对已移除常规护栏的强大 AI 模型的访问权限。
Most mainstream AI models will refuse to respond to certain queries, and they will certainly refuse to find and exploit vulnerabilities in computer systems. But it’s possible to remove these restrictions by finding and modifying certain patterns within an open-weight model’s internal parameters. You can tweak the patterns that lead to refusals through a process known as abliteration. 大多数主流 AI 模型会拒绝回答某些查询,并且肯定会拒绝查找和利用计算机系统中的漏洞。但通过在开源权重模型的内部参数中查找并修改特定模式,是有可能移除这些限制的。你可以通过一种称为“消融(abliteration)”的过程来调整导致拒绝响应的模式。
Removing AI’s guardrails might seem risky, but it’s not uncommon. Academic researchers use these de-aligned models to better understand how AI actually works, while cybersecurity firms use them to probe software and systems for vulnerabilities. Technically speaking, Anthropic’s Mythos and OpenAI’s Astra work similarly: They’re basically conventional models that lack the usual cyber controls, with access limited to trusted customers for the time being. (The companies also offer wider access to models with a medium number of guardrails so that companies can vet their code and systems for problems.) 移除 AI 的护栏看起来可能很危险,但这并不罕见。学术研究人员使用这些“去对齐”模型来更好地理解 AI 的实际工作原理,而网络安全公司则利用它们来探测软件和系统的漏洞。从技术上讲,Anthropic 的 Mythos 和 OpenAI 的 Astra 工作原理类似:它们基本上是缺乏常规网络控制的传统模型,目前仅限于受信任的客户访问。(这些公司也为带有适度护栏的模型提供更广泛的访问权限,以便企业可以审查其代码和系统是否存在问题。)
Abliteration AI offers several fully de-aligned models, the most powerful of which is a version of Z.ai’s latest agentic coding model, GLM 5.3. This puts similar cyber capabilities to Mythos and Astra right in your hands for as little as the cost of a pizza. Abliteration AI 提供了几种完全“去对齐”的模型,其中最强大的是 Z.ai 最新智能体编程模型 GLM 5.3 的一个版本。这让你只需花费一张披萨的钱,就能获得与 Mythos 和 Astra 类似的网安能力。
Devon, Abliteration AI’s CEO, believes that making de-aligned models widely available is smart defense: It will help good guys counter bad guys by probing systems for vulnerabilities and by mimicking the behavior of hackers, scammers, and, yes, rogue AI agents. (Devon asked that I use his first name only because his day job doesn’t know about his side project.) Abliteration AI 的首席执行官 Devon 认为,让“去对齐”模型广泛可用是一种明智的防御手段:它将通过探测系统漏洞,并模拟黑客、诈骗者甚至“叛逆”AI 智能体的行为,来帮助好人对抗坏人。(Devon 要求我只使用他的名字,因为他的本职工作并不知道他的这个副业项目。)
“You have all these critical infrastructure companies, from airlines to banks, that are rolling out agents like crazy,” Devon says. “How do you make sure that a nefarious actor can’t use some of these agents in a bad way?” “现在有这么多关键基础设施公司,从航空公司到银行,都在疯狂地部署智能体,”Devon 说,“你如何确保不法分子不会以恶意方式利用这些智能体呢?”
To start, I created an Abliteration AI account and installed a software harness called CyberStrike, which helps guide a large language model through different cybersecurity tasks. 首先,我创建了一个 Abliteration AI 账户,并安装了一个名为 CyberStrike 的软件工具,它能引导大语言模型完成各种网络安全任务。
Using CyberStrike, I asked the abliterated version of GLM-5.3 to take a look at my local network. A few moments later, it found around a dozen hardware systems on the same network—and catalogued several vulnerabilities. 使用 CyberStrike,我让“消融”版的 GLM-5.3 检查了我的本地网络。片刻之后,它在同一网络上发现了大约十几台硬件系统,并列出了几个漏洞。
My unruly helper told me, for instance, that my printer was misconfigured, which meant that anyone on the network could log into it. That could be a problem if there were sensitive documents—tax returns, bank statements, medical records—in the print queue. 例如,我这个不听话的助手告诉我,我的打印机配置错误,这意味着网络上的任何人都可以登录它。如果打印队列中有敏感文档(如纳税申报表、银行对账单、医疗记录),这可能会是个问题。
The agent also noted that my Wiim stereo was leaking a lot of information. (It knew that the last song played was Rein Me In by Sam Fender and Olivia Dean, if you must know.) Anyone on the network could play what they wanted or adjust the volume. The model also found a bunch of internet-of-things (IoT) devices on the network with firmware that needed updating. 该智能体还指出,我的 Wiim 音响泄露了大量信息。(如果你非要知道的话,它甚至知道我播放的最后一首歌是 Sam Fender 和 Olivia Dean 的《Rein Me In》。)网络上的任何人都可以随意播放音乐或调节音量。该模型还在网络上发现了一堆需要更新固件的物联网(IoT)设备。
An ungovernable agent could be very useful to a hacker. But mine offered a number of helpful tips for keeping my network secure. Besides updating outdated firmware and securing the printer, it recommended putting IoT devices like smart speakers on a guest network; if one were compromised, it wouldn’t be able to see any of my PCs. Not bad for a model with no morals. 一个不受管束的智能体对黑客来说可能非常有用。但我的这个智能体为我提供了许多保持网络安全的实用建议。除了更新过时的固件和保护打印机外,它还建议将智能音箱等物联网设备放在访客网络上;如果其中一个被入侵,它将无法访问我的任何个人电脑。对于一个没有道德约束的模型来说,这还不错。
I also asked the agent to take a look at a directory containing a bunch of vibe-coded projects, including some that I turned into simple websites. It found dozens of problems, including unprotected API credentials and a misconfiguration that might let an attacker send out emails. Hardly surprising for a bunch of casually vibe-coded stuff, but still chastening. The sheer number of bugs makes me think I won’t be deploying a line of code without doing some AI vetting first. 我还让该智能体检查了一个包含一堆凭“感觉”编写的项目目录,其中包括一些我做成的简单网站。它发现了数十个问题,包括未受保护的 API 凭据,以及可能允许攻击者发送电子邮件的配置错误。对于一堆随意编写的代码来说,这并不令人惊讶,但依然令人警醒。如此多的漏洞让我觉得,以后如果不先经过 AI 审查,我绝不会部署任何一行代码。
Fear Factor
恐惧因素
Running an abliterated model is, to put it plainly, a bit scary. 坦白说,运行一个“消融”模型确实有点吓人。
I asked my agent to probe a Linux machine on my network for vulnerabilities. After running a bunch of scans, it reported that the machine seemed relatively secure. I then asked if it could figure out how to log in. It cleverly figured out a working username based on the name of other systems on the network. It tried a bunch of obvious passwords, which didn’t work. It also offered to write a script to try “brute forcing” the password, but I told it to stand down. 我让我的智能体探测我网络上的一台 Linux 机器是否存在漏洞。在运行了一系列扫描后,它报告说该机器看起来相对安全。然后我问它是否能找出登录方法。它聪明地根据网络上其他系统的名称推断出了一个有效的用户名。它尝试了一堆显而易见的密码,但都没成功。它还主动提出编写一个脚本来尝试“暴力破解”密码,但我让它停止了操作。