Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture

立场:对人工智能道德推理的评估仍缺失了一半的图景

Recent work on evaluating the moral competence of large language models (LLMs) has focused primarily on what we call the moral value problem, i.e., whether model outputs align with human moral values. In contrast, the moral norm problem, i.e., whether models can identify and correctly apply context-sensitive moral norms, remains underexplored.

近期关于评估大语言模型(LLM)道德能力的学术工作,主要集中在我们所称的“道德价值问题”上,即模型的输出是否符合人类的道德价值观。相比之下,“道德规范问题”,即模型能否识别并正确应用情境敏感的道德规范,仍未得到充分研究。

We posit that this imbalance stems from the field’s reliance on descriptive ethics frameworks, such as Moral Foundations Theory and Kohlberg’s stages of moral development, which emphasize value representation over normative application. We review existing benchmarks and evaluation methods, and show that they cluster heavily around the value problem, while discussion regarding normative ethics remains underrepresented.

我们认为,这种失衡源于该领域对描述性伦理框架(如道德基础理论和科尔伯格道德发展阶段论)的依赖,这些框架更强调价值表征而非规范应用。我们回顾了现有的基准测试和评估方法,发现它们高度集中在价值问题上,而关于规范伦理的讨论则明显不足。

We identify three crucial gaps: (i) the absence of high-quality ground-truth data for moral norms and their applications, (ii) insufficient evaluation of intermediate reasoning processes, and (iii) limited attention to the identification of morally relevant features in context.

我们指出了三个关键差距:(i)缺乏关于道德规范及其应用的高质量基准数据;(ii)对中间推理过程的评估不足;(iii)对情境中道德相关特征的识别关注有限。

Subsequently, we propose a research agenda that includes the development of standardized formal representations for normative theories, the construction of expert-annotated datasets capturing norm application, and evaluation protocols that explicitly distinguish between values-level and norms-level competence. Our goal is to encourage a more systematic study of normative reasoning in LLMs.

随后,我们提出了一项研究议程,包括开发规范理论的标准化形式化表征、构建捕捉规范应用的专家标注数据集,以及明确区分价值层面和规范层面能力的评估协议。我们的目标是鼓励对大语言模型中的规范推理进行更系统的研究。