AI hallucination of Chinese nuclear components almost led to US military attack
AI hallucination of Chinese nuclear components almost led to US military attack
AI 对中国核部件的“幻觉”险些引发美军军事打击
The US narrowly avoided boarding a Chinese ship based on an “entirely false” US intelligence report generated with the help of AI tools, according to a CNN report. That erroneous intelligence, submitted by a US Special Operations Command analyst, suggested the Chinese ship was transporting nuclear arms program components through the Middle East, according to “four sources familiar with the episode” cited by CNN. The US military was preparing to intercept and board the ship, with air support, before officials discovered a chatbot used in generating the report had “inaccurately identified the material the ship was carrying.” One source told CNN the AI-powered fiasco “almost started a war.” 据 CNN 报道,美国险些因一份由 AI 工具生成的“完全虚假”的情报报告而登上一艘中国船只。据 CNN 引用的“四位知情人士”透露,这份由美国特种作战司令部分析师提交的错误情报称,该中国船只正在中东地区运输核武器计划的部件。在官员们发现用于生成该报告的聊天机器人“错误地识别了船上运载的物资”之前,美军已准备好在空中支援下拦截并登船。一位消息人士告诉 CNN,这场由 AI 引发的闹剧“差点引发了一场战争”。
What’s the worst that can happen? CNN’s report said the analyst in question used a chatbot to analyze intelligence reports regarding the Chinese ship’s manifest, leading to the near-disastrous result. That chatbot “fused together open-source intelligence with secret signals intelligence in government holdings,” and that information was packaged into an intelligence report that almost set off a disastrous chain of events, according to CNN. 最坏的情况会是什么?CNN 的报道称,涉事分析师使用聊天机器人分析了有关该中国船只货单的情报报告,从而导致了这一近乎灾难性的结果。据 CNN 称,该聊天机器人“将开源情报与政府持有的秘密信号情报融合在一起”,并将这些信息打包成一份情报报告,险些引发一连串灾难性事件。
The near-miss is one of the more potent and consequential instances of a hallucinating AI ruining the reliability of a professional report. Since “hallucinating” became the Cambridge Dictionary’s word of the year in 2023, we’ve seen prominent examples of non-fiction authors, journalists, academic researchers, judges, doctors, police departments, corporate call centers, and more getting taken in by AI tools that simply make something up when their training data doesn’t provide sufficient context. And despite some adorable attempts at “do not hallucinate” prompts, some researchers suggest that it may be impossible to prevent LLMs from hallucinating altogether. 这次险情是 AI “幻觉”破坏专业报告可靠性的最有力、后果最严重的案例之一。自“幻觉”(hallucinating)成为 2023 年剑桥词典年度词汇以来,我们已经看到许多非虚构类作家、记者、学术研究人员、法官、医生、警察部门、企业呼叫中心等被 AI 工具误导的典型案例——当训练数据无法提供足够的背景信息时,这些工具往往会凭空捏造。尽管人们尝试了一些可爱的“禁止幻觉”提示词,但一些研究人员认为,要完全防止大语言模型(LLM)产生幻觉可能是不可能的。
One would hope the US military would be aware of these kinds of problems when relying on AI for analysis of intelligence reports. But the Department of Defense in January rolled out an “AI acceleration strategy” that sought to “make all appropriate data available across federated IT systems for AI exploitation, including mission systems across every service and component.” “AI is only as good as the data that it receives, and we’re going to make sure that it’s there,” Defense Secretary Pete Hegseth said in rolling out that initiative. 人们本以为美军在依赖 AI 分析情报报告时会意识到这些问题。然而,美国国防部在今年 1 月推出了“AI 加速战略”,旨在“使所有适当的数据在联邦 IT 系统中可用,以供 AI 开发利用,包括各军种和部门的任务系统”。国防部长皮特·赫格塞斯(Pete Hegseth)在推出该倡议时表示:“AI 的水平取决于它所接收的数据,我们将确保这些数据到位。”
Last December, the Department of Defense announced it would use Google’s Gemini for Government as the basis for its bespoke “GenAI.mil” platform. Last month, the department added Grok for Government as an option on the platform. Anthropic also offers a customized version of Claude for US spy work. In June, a Pentagon representative bragged to Congress that they use generative AI to help create congressionally mandated reports, and that 1.5 million active DoD personnel have used the military’s generative AI tools. 去年 12 月,美国国防部宣布将使用谷歌的“政府版 Gemini”(Gemini for Government)作为其定制平台“GenAI.mil”的基础。上个月,该部门又将“政府版 Grok”(Grok for Government)添加为该平台的选项。Anthropic 公司也为美国情报工作提供了定制版的 Claude。今年 6 月,五角大楼的一位代表向国会吹嘘称,他们利用生成式 AI 来协助撰写国会要求的报告,并且已有 150 万名现役国防部人员使用了军方的生成式 AI 工具。
In 2023, a State Department “Declaration on Responsible Military Use of Artificial Intelligence and Autonomy” stressed that “principled” use of AI by armed forces “should include careful consideration of risks and benefits, and it should also minimize unintended bias and accidents.” That report also urged that “accountable” use of AI systems must always involve “a human in the loop, a responsible human chain of command and control.” In the years since then, though, we’ve seen fully autonomous attack drones used in the Russian conflict in Ukraine and tested by NATO-backed military contractors. 2023 年,美国国务院发布的《关于负责任地在军事中使用人工智能和自主权的声明》强调,武装部队对 AI 的“原则性”使用“应包括对风险和收益的仔细考量,并应最大限度地减少意外偏见和事故”。该报告还敦促,对 AI 系统的“负责任”使用必须始终包含“人在回路(human in the loop),以及负责任的人类指挥和控制链”。然而,在此后的几年里,我们已经看到全自动攻击无人机被用于俄乌冲突,并由北约支持的军事承包商进行测试。
In March, the Department of Defense blacklisted Anthropic over the company’s opposition to its models’ use in autonomous weapons systems, a move that a federal judge said last month was “unlawful retaliation in violation of the First Amendment.” The reported near miss comes as extinction-level warnings from AI researchers have led to a newly prominent national conversation on AI safety, including calls for regulation and coordinated research “pacing” from leading frontier AI labs. 今年 3 月,国防部将 Anthropic 列入黑名单,原因是该公司反对将其模型用于自主武器系统。一位联邦法官上个月表示,此举是“违反第一修正案的非法报复”。此次报道的险情发生之际,AI 研究人员发出的“灭绝级”警告引发了全国范围内关于 AI 安全的新一轮重要讨论,其中包括对监管的呼吁,以及要求领先的前沿 AI 实验室进行协调一致的研究“节奏控制”。