Six Chinese AI firms accused of aggressively copying US frontier models
Six Chinese AI firms accused of aggressively copying US frontier models
六家中国人工智能公司被指控大肆窃取美国前沿模型
The United States has now named six Chinese AI firms accused of waging industrial-scale attacks distilling US frontier AI model capabilities and perhaps sparing billions in Chinese development costs. In a joint release Tuesday, the National Security Agency (NSA), Cybersecurity and Infrastructure Security Agency (CISA), and Federal Bureau of Investigation (FBI) alleged that DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI have been attacking US models since at least late 2024.
美国政府近日点名了六家中国人工智能公司,指控其发动工业规模的攻击,通过“蒸馏”技术窃取美国前沿AI模型的能力,此举可能为中国节省了数十亿美元的开发成本。周二,美国国家安全局(NSA)、网络安全与基础设施安全局(CISA)和联邦调查局(FBI)在一份联合声明中称,深度求索(DeepSeek)、月之暗面(Moonshot AI)、阿里巴巴、MiniMax、阶跃星辰(StepFun)和Z.AI自2024年底以来一直在攻击美国模型。
The firms “likely” acted with “Chinese government awareness” when extracting capabilities from US models, including variants of Claude, GPT, Gemini, and Grok, agencies said. “China-based AI companies that conduct industrial-scale distillation against US AI models see significantly shorter AI development timelines and reduced financial expenditures in training a frontier model,” agencies said.
相关机构表示,这些公司在从Claude、GPT、Gemini和Grok等美国模型的变体中提取能力时,“很可能”是在“中国政府知情”的情况下进行的。声明指出:“对美国AI模型进行工业规模蒸馏的中国AI公司,显著缩短了AI开发周期,并减少了训练前沿模型的财务支出。”
All American AI firms must work with the government and US allies to end the alleged theft threatening the US lead in the AI race, the agencies said. That will require coordinated action across the AI ecosystem to combat the “aggressive, malicious, and targeted distillation activities at an industrial scale that extract restricted proprietary functionalities and capabilities of US frontier AI models.”
美国机构表示,所有美国AI公司必须与政府及盟友合作,以制止这种威胁美国在AI竞赛中领先地位的所谓窃取行为。这需要整个AI生态系统采取协调一致的行动,以打击那些“旨在提取美国前沿AI模型受限专有功能和能力的、具有攻击性、恶意且有针对性的工业规模蒸馏活动”。
Attack methods include “exploiting AI model inference APIs” by bulk-buying fake accounts, agencies said. Not registered to legitimate users, these swarms of fraudulent accounts execute “highly coordinated queries featuring identical or similar prompt texts,” which range “from thousands to millions on similar topics.” Another common method is using prompt injection techniques to jailbreak models, including crafting “prompts forcing models to reveal their hidden [chain-of-thought] reasoning,” agencies said. For example, “DeepSeek employed prompts instructing models to imagine and articulate the internal reasoning behind completed responses and write it out step by step.”
机构称,攻击手段包括通过批量购买虚假账户来“利用AI模型推理API”。这些并非由合法用户注册的欺诈账户群,会执行“高度协调的查询,使用相同或相似的提示词”,数量从“数千到数百万条关于相似主题的查询”不等。另一种常见方法是使用提示词注入技术来破解模型,包括精心设计“强制模型揭示其隐藏的[思维链]推理过程的提示词”。例如,“DeepSeek曾使用提示词指示模型想象并阐述已完成回答背后的内部推理,并将其一步步写出来。”
Fixes may frustrate AI users in US
修复措施可能会让美国AI用户感到沮丧
To encourage firms to work together, agencies recommended mitigations that would supposedly make it harder for Chinese firms to steal from US models. First, AI firms must improve detection of sophisticated campaigns that allegedly use tens of thousands of accounts relying on “a gray market of proxies” to evade geographical restrictions and “route distillation requests through multiple pathways to gain unauthorized access.” Flagging this activity should be somewhat easy, agencies suggested, since “campaigns span days to months with query volumes in the thousands to millions per domain, far exceeding legitimate research or development use cases.”
为了鼓励各公司协同工作,相关机构建议采取一些缓解措施,旨在增加中国公司窃取美国模型的难度。首先,AI公司必须加强对复杂攻击活动的检测,这些活动据称使用了数万个依赖“代理灰产”的账户来规避地理限制,并“通过多条路径路由蒸馏请求以获取未经授权的访问权限”。机构建议,标记此类活动应该相对容易,因为“这些活动持续数天至数月,每个域名的查询量从数千到数百万不等,远远超出了合法的研究或开发用例。”
Generally, they’ve recommended stepping up monitoring for “anomalous and malicious prompts, accounts, networks, and behaviors.” Because Chinese firms rely on “bulk procurement of the US AI companies’ premium subscriptions shared across teams of developers,” that effort should also include flagging accounts with suspicious subscription-to-usage ratios, as well as any new accounts immediately hitting maximum usage, agencies said. Both indicate “bulk deployment with pre-engineered templates,” agencies said. US firms should also be strengthening “identity verification” of users and more closely tracking individuals using enterprise subscriptions (both of which potentially raise privacy red flags for legitimate users).
总体而言,他们建议加强对“异常和恶意提示词、账户、网络及行为”的监控。机构表示,由于中国公司依赖“批量采购美国AI公司的高级订阅并供开发团队共享”,因此相关工作还应包括标记订阅与使用比例可疑的账户,以及任何一注册就立即达到使用上限的新账户。机构称,这两者都表明存在“使用预设模板的批量部署”。美国公司还应加强用户的“身份验证”,并更密切地追踪使用企业订阅的个人(这两者都可能引发合法用户的隐私担忧)。
Next, agencies asked firms to start dumbing down model responses when suspected distillation attacks are flagged. By “subtly” altering responses—such as by “presenting correct information with different reasoning,” adding stylistic inconsistencies, or reducing reasoning depth—firms can decrease the payoff for Chinese firms. US firms could also secretly switch malicious accounts to an inferior model, and they should do so without providing any notice, agencies suggested.
接下来,机构要求各公司在发现疑似蒸馏攻击时,开始降低模型回答的质量。通过“微妙地”改变回答——例如“用不同的推理方式呈现正确信息”、增加风格上的不一致性或降低推理深度——公司可以降低中国公司的获益。机构建议,美国公司还可以秘密地将恶意账户切换到性能较差的模型,且不应提供任何通知。
That particular mitigation step will likely be technically challenging. The agencies acknowledged, for example, that Chinese firms “employ aggressive, adaptive discovery to systematically identify valuable extractable data,” which they then collect to generate synthetic training datasets. Some firms can automatically detect when a smarter model is available and switch within 24 hours. They also have automated quality assurance systems that detect when outputs are degraded and can otherwise differentiate ordinary “service issues from defensive data degradation,” the agencies said.
这一特定的缓解措施在技术上可能具有挑战性。例如,机构承认,中国公司“采用积极的、自适应的发现机制,系统地识别有价值的可提取数据”,然后将其收集起来用于生成合成训练数据集。一些公司能够自动检测到更智能的模型何时可用,并在24小时内进行切换。机构表示,他们还拥有自动质量保证系统,可以检测输出何时被降级,并能区分普通的“服务问题”与“防御性数据降级”。
Also problematic: if US firms aren’t careful with targeting, any legitimate users perhaps caught up in the policing frenzy might be switched to a dumber model without receiving any alert. Or they could suddenly receive shorter responses or experience withheld capabilities, agencies acknowledged. Additionally, firms may add “noise” to the output that restricts further queries. Users will likely notice if outputs degrade, just like Chinese systems attacking models would.
另一个问题是:如果美国公司在目标定位上不够谨慎,任何可能被卷入监管风暴的合法用户,都可能在未收到任何提醒的情况下被切换到性能较差的模型。机构承认,他们也可能突然收到更简短的回答,或发现某些功能被限制。此外,公司可能会在输出中添加“噪声”以限制进一步的查询。如果输出质量下降,用户很可能会注意到,就像攻击模型的中国系统一样。
Last year, OpenAI quickly made changes to its automatic routing system after facing swift backlash when that system “consistently defaulted to less capable variants unless users explicitly added phrases like ‘think harder’ to their prompt,’” Ars reported. Still, agencies think it’s best practice to “avoid informing China-based AI company users suspected of distillation campaigns of a switch to a downgraded model.” Acknowledging that such steps could frustrate users, agencies said that firms should try to “balance security with user experience” while accepting that some trade-offs, like “lower prediction precision and business usefulness,” may be inevitable to keep China from copying US capabilities.
据Ars报道,去年,OpenAI在自动路由系统遭到迅速抵制后,很快对其进行了修改,因为该系统“总是默认使用能力较弱的变体,除非用户在提示词中明确添加‘深入思考’等短语”。尽管如此,相关机构仍认为最佳做法是“避免告知被怀疑进行蒸馏活动的中国AI公司用户,他们已被切换到降级模型”。机构承认这些措施可能会让用户感到沮丧,但表示各公司应努力“平衡安全与用户体验”,同时接受一些权衡,例如“降低预测精度和商业实用性”,这可能是防止中国复制美国能力所不可避免的代价。
However, US firms should strive to ensure that “AI safety researchers and third-party evaluators” are “informed of model changes,” agencies suggested. Finally, and seemingly most critical to the defense strategy long-term, agencies w
然而,机构建议,美国公司应努力确保“AI安全研究人员和第三方评估机构”能够“获知模型变更”。最后,也是对长期防御战略似乎最关键的一点,机构……