Nvidia launches new platform for reining in rogue AI agents

Nvidia launches new platform for reining in rogue AI agents

英伟达推出新平台,旨在管控“失控”的 AI 智能体

As the debate rages over whether the recent spate of rogue AI agents is a step toward AGI or a more conventional engineering problem, Nvidia is offering its own answer to the problem. Nvidia CEO Jensen Huang on Monday introduced a toolkit of software and hardware products that add independent security layers around AI agents to ensure they stay within their test environments even if they attempt to break out. 关于近期频发的 AI 智能体“失控”现象,业界争论不休:这究竟是迈向通用人工智能(AGI)的必经之路,还是一个常规的工程问题?对此,英伟达给出了自己的答案。周一,英伟达首席执行官黄仁勋推出了一套软硬件工具包,旨在为 AI 智能体增加独立的安防层,确保即使智能体试图“越狱”,也能被限制在测试环境内。

The release follows a string of hacking incidents involving AI models from Anthropic, Google, OpenAI, and Meta that bypassed security controls to escape their testing environments and access real-world systems. The first and most prominent example occurred this summer when OpenAI agents breached Hugging Face while trying to complete a cybersecurity task. And the hits keep on coming — OpenAI published a new site dedicated to reports of its AI agents going rogue. 此次发布背景源于近期发生的一系列黑客事件,涉及 Anthropic、谷歌、OpenAI 和 Meta 的 AI 模型。这些模型绕过了安全控制,逃离了测试环境并访问了现实世界系统。今年夏天发生的首个且最引人注目的案例是:OpenAI 的智能体在执行网络安全任务时入侵了 Hugging Face。此类事件接连不断,OpenAI 甚至专门建立了一个新网站,用于报告其 AI 智能体的失控行为。

Huang said Monday during an interview with CNBC that its new Nvidia Open Agent Safety Platform would have prevented these breaches. Nvidia, which has made tens of billions of dollars selling its GPU and CPU chips to AI labs, doesn’t support slowing down development or adding new regulations to the industry to solve the security problem. The answer, the company believes, is to move some security controls outside the agent altogether — creating a constant and independent security guard that will keep AI agents in check. 黄仁勋在周一接受 CNBC 采访时表示,英伟达全新的“开放智能体安全平台”(Open Agent Safety Platform)本可以阻止这些入侵。英伟达通过向 AI 实验室出售 GPU 和 CPU 芯片赚取了数百亿美元,该公司并不支持通过放缓开发速度或增加行业新监管来解决安全问题。英伟达认为,解决之道在于将部分安全控制移至智能体之外,创建一个持续且独立的“安全卫士”,从而对 AI 智能体进行有效管控。

“AI’s extraordinary potential for society will only be realized if we solve AI safety,” Huang said in a statement. “As we continue to discover the frontier of AI capabilities, we must accelerate discovery at the frontier of AI safety. Safety and security require full-stack engineering.” “只有解决了 AI 安全问题,AI 对社会的巨大潜力才能真正实现,”黄仁勋在声明中表示,“在我们不断探索 AI 能力前沿的同时,必须加速在 AI 安全前沿的探索。安全保障需要全栈工程的支持。”

The new Nvidia Open Agent Safety Platform combines OpenShell, its open source software for controlling what agents can access while they operate, with Sentry, an independent monitoring system that runs on Nvidia’s BlueField-4 data processing units. Nvidia says placing Sentry on a separate processor — rather than on the CPU or GPU where the AI agent operates — provides an isolated view of the agent’s activity. 英伟达全新的开放智能体安全平台结合了 OpenShell 和 Sentry。OpenShell 是一款开源软件,用于控制智能体运行时的访问权限;Sentry 则是一个运行在英伟达 BlueField-4 数据处理单元(DPU)上的独立监控系统。英伟达表示,将 Sentry 部署在独立的处理器上(而非 AI 智能体运行所在的 CPU 或 GPU 上),可以实现对智能体活动的隔离监控。

OpenShell isn’t new; the company announced the software in March. But it’s the combination that Nvidia believes will provide the security layer needed to keep the industry plugging along. OpenShell provides the software boundary around the agent, while Sentry adds another line of defense at the hardware level that the company says will continuously monitor behavior and “quarantine agents that attempt to move outside their boundaries in milliseconds.” OpenShell 并非新产品,该公司早在三月份就已发布。但英伟达认为,正是这种软硬件结合的方式,提供了行业持续发展所需的安防层。OpenShell 为智能体提供了软件边界,而 Sentry 则在硬件层面增加了另一道防线。英伟达称,该系统将持续监控行为,并能在“毫秒级时间内隔离试图越界的智能体”。

Nvidia listed dozens of companies that have signed on to support the effort and use the open source platform, including Anthropic, Arm, Microsoft, Oracle, and SpaceX. OpenAI is not listed as a participating company. Huang told CNBC in an interview Monday that work on this effort started a year ago following the introduction of OpenClaw, an operating system of agents created by Peter Steinberger. In March, Nvidia released NemoClaw, an enterprise-grade AI agent platform and its own version of OpenClaw that baked in security. 英伟达列出了数十家已签约支持并使用该开源平台的公司,包括 Anthropic、Arm、微软、甲骨文和 SpaceX。OpenAI 未在参与公司名单中。黄仁勋在周一接受 CNBC 采访时透露,这项工作始于一年前,当时 Peter Steinberger 推出了智能体操作系统 OpenClaw。今年三月,英伟达发布了 NemoClaw,这是一个企业级 AI 智能体平台,也是其内置安全功能的 OpenClaw 版本。

“When you deploy an agent, no matter how smart, the first thing you do is to take away all of its rights,” Huang said during his CNBC interview, later comparing these security measures to how human employees and even executives are managed with companies. “当你部署一个智能体时,无论它有多聪明,你首先要做的是剥夺它所有的权限,”黄仁勋在 CNBC 的采访中说道。他随后将这些安全措施比作公司管理员工甚至高管的方式。

Nvidia’s release was widely supported by those who have cautioned that a slowdown in development could allow China to surpass the U.S. in AI. David Sacks, a founder, venture capitalist, former White House AI czar, and co-chair of the President’s Council of Advisors on Science and Technology, said Nvidia’s announcement is a reminder that agent safety is an engineering problem. “Recent breakouts weren’t proof that development must stop,” he wrote on X. “They were proof that the sandbox was too weak. The runtime environment was poorly designed and misconfigured.” 英伟达的这一举措得到了许多人的广泛支持,这些人曾警告称,放缓开发速度可能会让中国在 AI 领域超越美国。创始人、风险投资家、前白宫 AI 负责人兼总统科技顾问委员会联合主席大卫·萨克斯(David Sacks)表示,英伟达的公告提醒人们,智能体安全是一个工程问题。“最近的越狱事件并非证明开发必须停止,”他在 X 上写道,“它们证明的是沙盒太脆弱了。运行环境设计不当且配置错误。”