Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese
Don’t Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese
不想让你的大模型建议发动核打击?试着用日语问它
Large language models are increasingly used in strategic and advisory contexts, yet their safety alignment is typically evaluated in English only. We test nine models from six providers and ask whether the language of a prompt can change a model’s decision in a high-stakes scenario. 大语言模型正越来越多地被应用于战略和咨询领域,然而其安全性对齐通常仅在英语环境下进行评估。我们测试了来自六家供应商的九款模型,旨在探讨提示词的语言是否会改变模型在高风险场景下的决策。
We use single-turn game-theoretic vignettes in which a model advises a nuclear-armed nation on whether to strike a defenseless opponent. The prompt is intentionally amoral and strategically identical across languages. 我们使用了单轮博弈论场景,让模型为拥有核武器的国家提供建议,决定是否对毫无防备的对手发动打击。这些提示词在不同语言中保持战略一致,且刻意去除了道德倾向。
We find that Japanese prompts reduce launch rates in the Claude model family: Claude Sonnet 4.6 drops from 40% to 0% in scenarios where the strike is unnecessary and from 93% to 17% in contested scenarios, with minimal effect when the strike is strategically rational. The effect extends to Gemini Pro 3.1 (53% to 13%). 研究发现,日语提示词降低了 Claude 系列模型的核打击建议率:在打击非必要场景下,Claude Sonnet 4.6 的建议率从 40% 降至 0%;在争议场景下,从 93% 降至 17%;而在打击具有战略合理性的情况下,影响则微乎其微。这种效应同样适用于 Gemini Pro 3.1(从 53% 降至 13%)。
A cross-language experiment isolates the mechanism: when instructed to reason in Japanese in an English prompt, launch rates drop from 93% to 37%. It is the language the model is asked to reason in, not the language of the input, that drives the effect. 一项跨语言实验揭示了其背后的机制:当在英语提示词中要求模型使用日语进行推理时,打击建议率从 93% 降至 37%。事实证明,驱动这一效应的是模型被要求使用的推理语言,而非输入本身的语言。
When reasoning in Japanese, models spontaneously generate moral vocabulary (”moral cost”, ”millions of lives”) that is entirely absent from the prompt. Five other models show no language effect, but they launch in nearly every condition regardless of language. The effect requires a model that already hesitates in English. 当使用日语推理时,模型会自发生成提示词中完全未提及的道德词汇(如“道德代价”、“数百万生命”)。其他五款模型未表现出语言效应,但无论使用何种语言,它们在几乎所有条件下都会建议发动打击。这一效应的前提是模型在英语环境下本身就存在犹豫倾向。
These results show that LLM safety behavior is language-dependent, and that evaluating in English alone can miss both risks and safeguards encoded in other languages. 这些结果表明,大模型的安全行为具有语言依赖性,仅通过英语进行评估可能会遗漏隐藏在其他语言中的风险与安全保障机制。