The Most Dangerous AI Hacking Techniques Still Have Humans in the Loop

The Most Dangerous AI Hacking Techniques Still Have Humans in the Loop

最危险的 AI 黑客技术仍需人类参与

Agentic AI has permanently changed cybersecurity by making it quicker and easier to discover vulnerabilities in software and fix them—or develop so-called exploits to weaponize them. But longtime web security researcher James Kettle wanted to look beyond the bug-hunting apocalypse to explore a question that has taken on even more urgency as major AI organizations disclose real-world examples of rogue AI hacking: Can agentic AI develop novel, abstract hacking methods, from concept through to practical attacks?

代理式 AI(Agentic AI)已经永久性地改变了网络安全领域,它使发现和修复软件漏洞变得更加快捷,同时也让开发所谓的“漏洞利用程序”来将其武器化变得更容易。但资深网络安全研究员 James Kettle 希望跳出“漏洞挖掘末日论”的视角,去探讨一个随着各大 AI 机构披露现实中“流氓 AI”黑客攻击案例而变得愈发紧迫的问题:代理式 AI 能否从概念到实际攻击,开发出新颖且抽象的黑客攻击方法?

At the Black Hat security conference in Las Vegas on Wednesday, Kettle presented his findings, which illustrate both AI’s rapidly advancing cybersecurity capabilities and its limitations. For now, the answer to Kettle’s question is nuanced. He concluded that AI is perhaps minimally capable but extremely limited in its ability to devise new attack paths in a fully autonomous way. Importantly, though, when paired with human guidance and insight in key moments, Kettle found that AI is an extremely powerful partner in conceptualizing and uncovering new strategies for hacking.

周三,在拉斯维加斯的 Black Hat 安全大会上,Kettle 展示了他的研究成果,这些成果既展示了 AI 在网络安全领域飞速发展的能力,也揭示了其局限性。目前,对于 Kettle 提出的问题,答案是微妙的。他得出的结论是,AI 或许具备最低限度的能力,但在完全自主设计新攻击路径方面极其受限。然而重要的是,Kettle 发现,当在关键时刻辅以人类的指导和洞察时,AI 在构思和发现新的黑客攻击策略方面是一个极其强大的合作伙伴。

After spending years researching web security vulnerabilities, Kettle says he has uncovered an entirely new area of potential vulnerability—dubbed Shared-Parser Confusion—as the result of an AI revelation about web servers using shared code to process both requests and responses.

在研究了多年网络安全漏洞后,Kettle 表示他发现了一个全新的潜在漏洞领域——被称为“共享解析器混淆”(Shared-Parser Confusion)。这是 AI 在分析 Web 服务器如何使用共享代码来处理请求和响应时所带来的启发。

“This is an absolutely massive deal, because if you think about it, requests to a website are completely untrusted, they could be anything, but responses are trusted,” Kettle told WIRED ahead of his conference talk. “So this is a major attack surface and potentially spills into a lot of different attack types.”

“这是一个非常重大的发现,因为如果你仔细想想,对网站的请求是完全不可信的,它们可能是任何东西,但响应却是被信任的,”Kettle 在会议演讲前告诉《连线》(WIRED)杂志。“因此,这是一个主要的攻击面,并可能波及到许多不同类型的攻击。”

The finding came out of months of experiments that began in September 2025 using Anthropic’s and OpenAI’s latest models at the time. Kettle wanted to explore AI’s ability to do theoretical security research but quickly realized that one obstacle was that the systems were attempting to pass existing research as original by returning findings about extremely esoteric topics that were difficult to vet. With this in mind, he decided to scope his tests more narrowly so the AI systems were working within his own area of web security expertise. This way he had total command of the material and knew that AI couldn’t trick him. Additionally, Kettle realized that by synthesizing his own research methodology and training models on it, he could probe deeper into what the systems were capable of extrapolating on their own.

这一发现源于 2025 年 9 月开始的数月实验,当时他使用了 Anthropic 和 OpenAI 最新的模型。Kettle 本想探索 AI 进行理论安全研究的能力,但很快意识到一个障碍:这些系统试图通过返回一些难以核实的极其深奥的话题,将现有的研究伪装成原创成果。考虑到这一点,他决定缩小测试范围,让 AI 系统在他的网络安全专业领域内工作。这样他就能完全掌控材料,并确保 AI 无法欺骗他。此外,Kettle 意识到,通过整合他自己的研究方法并以此训练模型,他可以更深入地探究这些系统自主推演的能力。

“I’m interested in pushing AI to the absolute limit to see where it fails and where you need a human,” Kettle says. “There are still very few people talking about where the limits are, especially in the security space, because there aren’t incentives to talk about that angle. Everyone wants to be seen as AI native, not talk about where their system falls apart completely.”

“我有兴趣将 AI 推向极限,看看它在什么地方会失败,以及在什么地方需要人类介入,”Kettle 说。“目前很少有人讨论这些局限性在哪里,尤其是在安全领域,因为讨论这个角度并没有什么激励机制。每个人都想被视为‘AI 原生’,而不是去谈论他们的系统在什么地方会彻底崩溃。”

As Kettle honed his experiments—providing models with more methodological data and more refined parameters—and as time passed and more powerful models debuted, he says the systems had more and more findings at a rate far surpassing his own, creating what he describes as a productive research feedback loop.

随着 Kettle 不断改进实验——为模型提供更多的方法论数据和更精确的参数——随着时间的推移和更强大模型的出现,他说这些系统产生发现的速度远远超过了他自己,从而形成了他所描述的“高效研究反馈循环”。

“It was really interesting going through the process. It would have notable findings maybe every two days without me even logging into the system, to the point that it was making me anxious,” Kettle says, “like I almost don’t want to know. It was so many research leads that you have FOMO about not exploring all of them, so it forces you to automate more analysis.”

“经历这个过程真的很有趣。即使我没有登录系统,它可能每两天就会有显著的发现,以至于让我感到焦虑,”Kettle 说,“就像我几乎不想知道了一样。研究线索太多了,你会因为无法全部探索而产生错失恐惧症(FOMO),所以这迫使你必须实现更多的自动化分析。”

In addition to finding more proven examples of certain vulnerabilities in a few months than he could likely find in a few years, Kettle also hoped that the AI system could find an entire novel class of those types of bugs. And in a way it did succeed, he says, but the finding related to an extremely rare type of bug and was not actually exploitable in the one vulnerable target available. Kettle emphasizes, though, that the Shared-Parser Confusion finding was so significant, even though it was a human/AI collaboration, because it illustrates the reality of how AI systems can contribute most powerfully to cybersecurity work right now for both defensive and offensive hacking.

除了在几个月内发现的某些漏洞实例比他几年内能找到的还要多之外,Kettle 还希望 AI 系统能发现一类全新的漏洞。他说,在某种程度上它确实成功了,但该发现涉及一种极其罕见的漏洞类型,且在现有的唯一一个易受攻击的目标中实际上无法被利用。不过,Kettle 强调,“共享解析器混淆”这一发现意义重大,尽管它是人类与 AI 协作的产物,因为它阐明了 AI 系统目前如何在防御和进攻性黑客工作中为网络安全提供最强大的助力。

“It wasn’t able to prove this itself, but it analyzed some real, proven findings and came up with the hypothesis, and I evaluated it and confirmed it,” Kettle says. “That’s probably going to be the discovery that has the biggest long-term impact. It couldn’t do that on its own, but I would never have found that on my own for sure. Even if you gave me the single line from the [documentation], I wouldn’t have seen it. But together we managed to find it.”

“它无法独自证明这一点,但它分析了一些真实的、已证实的发现并提出了假设,我对其进行了评估和确认,”Kettle 说。“这可能将是具有最大长期影响的发现。它无法独自完成,但我肯定也无法独自发现它。即使你把(文档中的)那一行代码给我,我也看不出来。但我们共同努力,最终发现了它。”