Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments
Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments
共识并非对齐:人类与大语言模型在道德判断上的基础分歧
Abstract: Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human annotators and models rely on the same moral grounds. Two agents may reach the same judgment while appealing to different principles, contextual assumptions, or interpretations of the situation.
摘要: 与人类判断达成共识是评估大语言模型(LLM)对齐程度的常用代理指标。然而,最终标签的一致性并不能证明人类标注者与模型依赖于相同的道德基础。两个主体可能在得出相同判断的同时,却诉诸于不同的原则、语境假设或对情境的解读。
We test this distinction using a curated 500-item ETHICS-derived benchmark spanning five domains of moral judgment, with new human annotator and LLM annotations of both final labels and supporting rationales. Across frontier and open model families, agreement with human annotator majority labels is often high.
我们利用一个精心策划的、源自 ETHICS 基准的 500 项测试集对这一区别进行了验证,该测试集涵盖了五个道德判断领域,并包含了人类标注者和 LLM 对最终标签及其支持性理由的新标注。在各类前沿模型和开源模型系列中,模型与人类标注者多数意见的一致性往往很高。
However, rationale-level analysis reveals systematic divergence in the moral grounds expressed by human annotators and models. In particular, models redistribute attention across categories such as harm, respect, promise-keeping, justice, desert, and excuse relevance, even when their final labels match the human annotator majority.
然而,在理由层面的分析揭示了人类标注者与模型在表达道德基础时存在系统性的分歧。特别是,即使模型最终给出的标签与人类标注者的多数意见相符,它们在伤害、尊重、守信、正义、应得性以及借口相关性等类别上的关注点分配也存在差异。
Our results show that agreement should not be treated as equivalent to alignment. Label-based evaluation can therefore be misleadingly reassuring unless complemented by analysis of the reasons, principles, and moral priorities expressed in model judgments.
我们的研究结果表明,不应将“共识”等同于“对齐”。因此,除非辅以对模型判断中所表达的理由、原则和道德优先级的分析,否则仅基于标签的评估可能会产生误导性的安全感。