Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models
Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models
超越对错:评估大语言模型中的二阶社会推理
Abstract: Previous AI alignment efforts have focused primarily on first-order social norms — teaching models what is socially acceptable or unacceptable (e.g., `do not steal’). However, social intelligence depends not only on norm recognition, but also on anticipating who will enforce it and how (e.g., public shame or even imprisonment). These second-order expectations, known as metanorms, govern how people respond when social rules are broken.
摘要: 此前的人工智能对齐工作主要集中于一阶社会规范——即教导模型什么是社会可接受或不可接受的行为(例如“不要偷窃”)。然而,社会智能不仅依赖于对规范的识别,还依赖于预判谁会执行规范以及如何执行(例如公开羞辱甚至监禁)。这些被称为“元规范”(metanorms)的二阶期望,决定了当社会规则被破坏时人们会如何反应。
We introduce a novel framework for evaluating metanorm reasoning in Large Language Models (LLMs) along two dimensions: emotional appraisal and behavioral response, and propose new classification tasks, namely, predicting self-regulation in violators, and other-regulation in observers. We release a multi-perspective dataset, NormReact, of 450 norm violation scenarios, hand-annotated for emotions and behavioral responses across norm violators’ gender and observers’ social closeness.
我们引入了一个用于评估大语言模型(LLM)中元规范推理的新框架,该框架涵盖两个维度:情感评估和行为反应。同时,我们提出了新的分类任务,即预测违规者的自我调节以及观察者的他人调节。我们发布了一个名为 NormReact 的多视角数据集,其中包含 450 个规范违规场景,并针对违规者的性别和观察者的社会亲密度,对情感和行为反应进行了人工标注。
Current LLMs portray a harsher social world: across six models, they overpredict negative sanctions where humans would expect inaction, and alignment with human judgments deteriorates as social distance increases. These findings suggest that AI systems in norm-sensitive domains from conflict mediation to policy simulation, may risk producing a distorted picture of social regulation: one that over-represents punishment and under-represents the tolerance, restraint, and relational calibration that characterize actual norm enforcement in real world.
当前的大语言模型描绘了一个更为严苛的社会:在六个模型中,当人类预期会保持沉默时,模型却过度预测了负面制裁;且随着社会距离的增加,模型与人类判断的一致性会下降。这些发现表明,在从冲突调解到政策模拟等对规范敏感的领域中,人工智能系统可能会产生扭曲的社会监管图景:即过度强调惩罚,而忽视了现实世界中规范执行所特有的宽容、克制和关系校准。