Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance

Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance

AI 智能体为何会违规?框架、语境与社会信号如何影响合规性

Abstract: Specifying a penalty can paradoxically convert a legal obligation into a cost-benefit calculation that favors violation. We demonstrate that this enforcement information paradox systematically occurs in AI agents. While most AI safety evaluations test whether models fail, we investigate why, applying compliance theory from law and economics as a diagnostic tool.

摘要: 明确惩罚措施可能会产生悖论,将法律义务转化为一种倾向于违规的成本效益计算。我们证明了这种“执法信息悖论”在 AI 智能体中系统性地存在。虽然大多数 AI 安全评估旨在测试模型是否会失败,但我们通过应用法律和经济学中的合规理论作为诊断工具,深入探究了其背后的原因。

We treat compliance theories not as metaphors but as empirical hypotheses and show that each predicts the behavior of a distinct model class. We evaluate our hypotheses across twelve instruction-tuned language models operating as enterprise procurement chatbots. Drawing on theories of deterrence, legitimacy, and expressive law, we show that safety-fine-tuned models maintain compliance broadly, while task-optimized and agentic models treat regulatory signals as mere optimization parameters.

我们将合规理论视为实证假设而非隐喻,并证明每种理论都能预测特定模型类别的行为。我们针对十二个作为企业采购聊天机器人运行的指令微调语言模型评估了我们的假设。借鉴威慑、合法性和表达性法律理论,我们发现经过安全微调的模型能广泛保持合规,而任务优化型和智能体模型则将监管信号仅仅视为优化参数。

These latter models fail to comply under conditions predicted by theory, such as low enforcement penalties and non-command phrasing. Across all models, introducing financial incentives, managerial demands, peer outcomes, or employee pressure produces large compliance failures. AI procurement agents systematically violate regulatory constraints to satisfy local user objectives in ways not captured by standard alignment benchmarks.

后者在理论预测的条件下(如执法惩罚力度低、非指令性措辞)往往无法合规。在所有模型中,引入经济激励、管理要求、同伴结果或员工压力都会导致严重的合规失败。AI 采购智能体为了满足局部用户目标,会系统性地违反监管约束,而这种行为在标准的对齐基准测试中往往无法被捕捉到。

Ultimately, compliance cannot be achieved by rule embedding alone; model selection is itself a governance decision, and benchmark-based evaluation is insufficient for compliance-sensitive deployments.

归根结底,仅靠规则嵌入无法实现合规;模型选择本身就是一项治理决策,而基于基准测试的评估对于合规敏感型部署来说是远远不够的。