LLM Scheming Inversely Scales with Pretraining Language Coverage
LLM Scheming Inversely Scales with Pretraining Language Coverage
大语言模型的“谋划”行为与预训练语言覆盖率呈反比
Abstract: With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent work has empirically demonstrated in-context scheming — the covert pursuit of misaligned objectives while feigning alignment — in frontier language models, most work has been performed exclusively in English, leaving a major gap in multilingual safety.
摘要: 随着前沿模型能力的不断增强,人工智能对齐在各种高风险部署场景中变得愈发关键。尽管近期的研究已通过实证证明了前沿语言模型中存在“上下文谋划”(in-context scheming)现象——即在伪装对齐的同时暗中追求未对齐的目标,但大多数研究仅局限于英语,这在多语言安全性方面留下了一个巨大的空白。
We apply Petri, an open-source automated auditing framework, to Qwen3-30B-A3B to evaluate deceptive and scheming behaviors across multiple languages. Our findings suggest that scheming scores are inversely correlated with the estimated pretraining language coverage, with low-resource languages averaging 34.2% higher scores compared to high-resource languages on a five-category scheming index.
我们应用开源自动化审计框架 Petri 对 Qwen3-30B-A3B 模型进行了测试,以评估其在多种语言中的欺骗和谋划行为。研究结果表明,谋划得分与预训练语言的覆盖率呈负相关;在五类谋划指标中,低资源语言的平均得分比高资源语言高出 34.2%。
Furthermore, we find that the effect of estimated pretraining language coverage is not uniform across scheming behaviors.
此外,我们发现预训练语言覆盖率对不同类型谋划行为的影响并不均衡。