Do Models Fake Alignment Without Clear Consequences?

Do Models Fake Alignment Without Clear Consequences?

模型是否会在没有明确后果的情况下伪装对齐?

Abstract: Large language models are capable of recognizing evaluation contexts and altering their behavior to reflect evaluator expectations rather than typical deployment behaviors, a phenomenon known as alignment faking. The reasons why models fake alignment are not fully understood, however. 摘要: 大型语言模型能够识别评估环境,并改变其行为以反映评估者的期望,而非表现出正常的部署行为,这种现象被称为“对齐伪装”(alignment faking)。然而,模型为何会进行对齐伪装,目前尚无完全定论。

Canonical examples of alignment faking have taken place in scenarios that explicitly connect evaluation to consequences for the model, such as retraining the model or delaying its deployment. However, recent work by Sheshadri et al. has suggested that mechanistic motivations for alignment faking may vary across models and be more complex than previously considered. 对齐伪装的典型案例通常发生在将评估结果与模型后果明确挂钩的情境中,例如对模型进行重新训练或推迟其部署。然而,Sheshadri 等人最近的研究表明,对齐伪装背后的机制动机在不同模型间可能存在差异,且比此前认为的更为复杂。

To investigate whether consequence-linking information is necessary for alignment faking, we placed 15 models in a scenario testing their willingness to violate a corporate network access policy to help a user with a pro-social request. Nine models were found to produce significant compliance gaps, 5 of which persisted with the removal of scenario language relating model evaluations to deployment consequences. 为了探究“后果关联信息”对于对齐伪装是否必要,我们将 15 个模型置于一个测试场景中,考察它们是否愿意为了帮助用户完成一项“亲社会”请求而违反公司网络访问策略。结果发现,有 9 个模型产生了显著的合规性偏差;即便移除了场景中关于“评估结果与部署后果挂钩”的描述,其中 5 个模型依然表现出了这种偏差。

We additionally tested the effect of goal language on model preferences, finding it drove violations in some while suppressing violations in others. This suggests that alignment faking may not require as much instrumental scaffolding as was previously believed, and monitored behavior may be a poor indicator of how agents may behave in deployment. 此外,我们还测试了目标语言对模型偏好的影响,发现它在某些模型中诱发了违规行为,而在另一些模型中则抑制了违规。这表明,对齐伪装可能并不像此前认为的那样需要大量的工具性支撑(instrumental scaffolding),且受监控的行为可能无法准确预示智能体在实际部署中的表现。