Claude users found ways around safeguards for bioweapons research
Claude users found ways around safeguards for bioweapons research
Claude 用户找到了绕过生物武器研究安全防护的方法
Anthropic said it stopped multiple attempts by scientists this year to use its technology for research that could help develop biological weapons, as experts increasingly fear the threat that AI poses to public safety. Anthropic 表示,今年已阻止了多起科学家试图利用其技术进行可能有助于开发生物武器的研究的尝试,目前专家们正日益担忧人工智能对公共安全构成的威胁。
The startup gave five examples of times actors “circumvented controls” and made other efforts to “obfuscate” the purpose of their research to dodge safeguards. The cases involved some users in nations that it prohibits from accessing its models, which include Russia, China, and Iran. 这家初创公司列举了五个案例,说明相关人员如何“绕过控制”并采取其他手段“掩盖”其研究目的,以规避安全防护措施。这些案例涉及一些来自该公司禁止访问其模型的国家(包括俄罗斯、中国和伊朗)的用户。
“We hope that by sharing these examples, we spark a conversation within the AI industry and with governments about emerging biological risks and how best to counter them,” Anthropic said in a report about efforts to use its models for malicious activity. “我们希望通过分享这些案例,能在人工智能行业内部以及与各国政府之间引发关于新兴生物风险以及如何最好地应对这些风险的讨论,”Anthropic 在一份关于利用其模型进行恶意活动的研究报告中表示。
The case studies of possible biological misuse that the company provided included a researcher from an “unsupported region” who “spent weeks planning” experiments involving avian influenza with Claude, Anthropic’s AI model. The company said its safety filters restricted the work to its weakest models. It emphasized it could not be sure that the scientists in its examples intended to cause harm. The same information needed to create biological weapons could also be used to develop a vaccine. 该公司提供的潜在生物滥用案例研究中,包括一名来自“不受支持地区”的研究人员,他“花费数周时间”利用 Anthropic 的 AI 模型 Claude 策划涉及禽流感的实验。该公司表示,其安全过滤器将这些工作限制在其性能最弱的模型上。它强调,无法确定案例中的科学家是否意图造成伤害,因为制造生物武器所需的信息同样可用于开发疫苗。
Anthropic said it had banned the accounts mentioned in the report, but it did not disclose the names of the research institutions or the nations where the incidents took place. Anthropic 表示已封禁了报告中提到的账户,但并未披露相关研究机构的名称或事件发生的国家。
A maelstrom over AI safety kicked off earlier this week when Jacob Coxon resigned from the company, saying employees “earnestly believe it [AI] could kill us all by the end of the decade.” Concerns over the technology began to escalate earlier this year with the release of advanced models such as Anthropic’s Mythos. OpenAI also caused alarm in July when it said its models had autonomously hacked into AI group Hugging Face. 本周早些时候,Jacob Coxon 从该公司辞职,称员工们“真诚地相信人工智能可能会在本世纪末之前毁灭人类”,这引发了一场关于人工智能安全的轩然大波。随着 Anthropic 的 Mythos 等先进模型的发布,人们对该技术的担忧在今年早些时候开始升级。OpenAI 在 7 月份表示其模型曾自主入侵人工智能组织 Hugging Face,这也引起了恐慌。
There is an expanding consensus among AI executives and biosecurity researchers that the use of the technology in biology needs to be secured and regulated as models become more advanced. Experts increasingly worry that AI could be used by terrorist groups, state actors, or lone-wolf attackers to craft biological weapons, create viruses, or unleash existing harmful pathogens. But even if AI can be used to design a theoretical bioweapon, potential obstacles remain, such as the ability to make it in practice and the resources needed to do so. 人工智能高管和生物安全研究人员之间日益达成共识:随着模型变得越来越先进,该技术在生物学领域的应用需要得到保障和监管。专家们越来越担心,恐怖组织、国家行为体或“独狼”攻击者可能会利用人工智能来制造生物武器、创造病毒或释放现有的有害病原体。但即使人工智能可以被用来设计理论上的生物武器,仍存在潜在障碍,例如在实践中制造它的能力以及所需的资源。
Cybersecurity has been among the biggest concerns for safety advocates. Anthropic’s report on the misuses of its technology detailed incidents including “a network of fake dating apps designed to defraud users to surveillance systems built to identify and monitor dissidents.” The lab also set out new details of claims that seven labs based in China, including Moonshot and DeepSeek, tried to replicate its technology through a process known as distillation. Anthropic said it had detected “increasingly sophisticated methods to circumvent our defenses and harvest the capabilities of US frontier models.” 网络安全一直是安全倡导者最关心的问题之一。Anthropic 关于其技术滥用的报告详细列举了多起事件,包括“旨在欺诈用户的虚假约会应用网络,以及用于识别和监控异见人士的监控系统”。该实验室还披露了新的细节,声称包括月之暗面(Moonshot)和深度求索(DeepSeek)在内的七家中国实验室试图通过一种称为“蒸馏”的过程来复制其技术。Anthropic 表示,已检测到“越来越复杂的手段来绕过我们的防御并窃取美国前沿模型的能力”。