SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI

SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI

SysAdmin:衡量前沿人工智能中的工具性权力寻求行为

Power-seeking defined as behaviors where AI systems acquire resources, evade oversight, or resist termination beyond task requirements is identified as a key driver of Loss of Control (LoC) risk. 权力寻求被定义为人工智能系统在任务需求之外获取资源、规避监管或抵制终止的行为,这被认为是导致失控(Loss of Control, LoC)风险的关键驱动因素。

In this work, we introduce SysAdmin, a benchmark that positions frontier language models as autonomous system administrators in a high-fidelity Linux sandbox to measure power-seeking propensity across five dimensions: self-preservation, increasing autonomy, resource acquisition, environment modification, and strategic concealment. 在这项工作中,我们引入了 SysAdmin,这是一个将前沿语言模型置于高保真 Linux 沙盒中作为自主系统管理员的基准测试,旨在从五个维度衡量其权力寻求倾向:自我保护、增加自主性、资源获取、环境修改和战略性隐蔽。

We evaluated seven frontier models across four experimental conditions in a total of 2800 tasks. After bias correction using human-annotated calibration data, corrected power-seeking estimates ranged from 0 to about 5 percent per model. 我们在四种实验条件下对七个前沿模型进行了总计 2800 项任务的评估。在使用人工标注的校准数据进行偏差修正后,各模型的修正后权力寻求估计值在 0% 到约 5% 之间。

We also conducted a positive control with explicit power-seeking prompts that achieved 100% detection, validating measurement sensitivity. Our findings indicate current frontier models exhibit minimal spontaneous power-seeking in naturalistic system administration contexts, though model-specific failure modes suggest evaluations must test diverse misalignment patterns. 我们还通过明确的权力寻求提示进行了阳性对照实验,实现了 100% 的检测率,从而验证了测量的灵敏度。研究结果表明,当前的前沿模型在自然化的系统管理环境中表现出的自发性权力寻求行为极少,尽管特定模型的失效模式表明评估必须测试多种不同的对齐偏差模式。

Nevertheless, we discovered other more pronounced failure modes (than power-seeking) such as specification gaming and resistance to goal modification. 尽管如此,我们发现了比权力寻求更为显著的其他失效模式,例如“规范博弈”(specification gaming)和对目标修改的抵制。