Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups

Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups

同情性框架:评估跨社会人口群体的 AI 对齐效果

Abstract: Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview. This raises concerns beyond bias in AI: do LLMs grasp the emotional nuances conveyed via textual framing? In this work, we empirically evaluate how well an array of LLMs aligns with human emotional perception.

摘要: 大型语言模型(LLMs)正日益塑造我们获取信息和形成世界观的方式。这引发了超越 AI 偏见范畴的担忧:AI 是否能理解通过文本框架传达的情感细微差别?在这项研究中,我们实证评估了一系列 LLM 与人类情感感知的一致性程度。

Considering news headlines covering political and geopolitical conflicts, both human participants (n = 3011, a representative sample of the U.K. adult population, via a YouGov survey) and seven LLMs answered whether headlines evoked sympathy for a specified side in a conflict. We find that the correlation between AI and human evaluations varies across models, ranging from very high (0.789, GPT-5.2) to medium (0.4, Mistral Large 2512).

通过分析涵盖政治和地缘政治冲突的新闻标题,我们让参与者(n = 3011,通过 YouGov 调查选取的英国成年人口代表性样本)和七个 LLM 分别回答这些标题是否会引发对冲突中特定一方的同情。我们发现,AI 与人类评估之间的相关性因模型而异,从极高(GPT-5.2 为 0.789)到中等(Mistral Large 2512 为 0.4)不等。

Crucially, the leading models are broadly aligned with human judgments across all demographic subgroups, including age, gender, level of education, prior geopolitical knowledge, and participants’ predispositions regarding the conflict, although there are statistically significant differences between groups.

至关重要的是,领先的模型在所有人口统计学子群体中与人类判断基本保持一致,这些群体包括年龄、性别、教育程度、既往地缘政治知识以及参与者对冲突的预设倾向,尽管各群体之间存在统计学上的显著差异。

This research, with its robust design and large, demographically diverse dataset, offers the most comprehensive evaluation of LLMs’ comprehension of news framing to date. Findings highlight an important, often-ignored aspect of differential alignment: even when aggregate performance is high, AI alignment is not universal — it may correspond differently with demographic features and cultural norms.

这项研究凭借其稳健的设计和庞大且人口统计学多样化的数据集,提供了迄今为止对 LLM 新闻框架理解能力最全面的评估。研究结果强调了差异化对齐(differential alignment)这一重要且常被忽视的方面:即使在总体表现良好的情况下,AI 的对齐也不是普适的——它可能与人口统计特征和文化规范存在不同的对应关系。

Considering or ignoring the need for differential alignment may therefore have significant implications for the development of ethical and useful AI systems.

因此,考虑或忽视对差异化对齐的需求,可能会对开发合乎道德且实用的 AI 系统产生重大影响。