Learning Sexism Detection Using Multi-Agent Perspectivist Preference Optimization
Learning Sexism Detection Using Multi-Agent Perspectivist Preference Optimization
使用多智能体视角偏好优化学习性别歧视检测
Abstract: When people label text for sexism, they often disagree, and not because some of them are wrong: they genuinely perceive sexism differently. Most NLP systems discard this disagreement by collapsing it into a majority vote. We propose the Multi-Agent Perspectivist Preference Optimization (MAP-PO) framework to keep these different perspectives.
摘要: 当人们对文本进行性别歧视标注时,往往会产生分歧,这并非因为其中某些人是错误的,而是因为他们对性别歧视的感知确实存在差异。大多数自然语言处理(NLP)系统通过将这些分歧归纳为多数投票来丢弃它们。我们提出了多智能体视角偏好优化(MAP-PO)框架,以保留这些不同的视角。
On the EXIST 2024 dataset of labeled English and Spanish tweets, we first cluster annotators by their labeling behavior rather than their demographic attributes. We then fine-tune one Large Language Model agent per cluster to reproduce that cluster’s annotation behavior, and coordinate the agents with preference optimization that combines individual and team-level rewards.
在包含已标注英语和西班牙语推文的 EXIST 2024 数据集上,我们首先根据标注者的标注行为而非人口统计学属性对他们进行聚类。随后,我们为每个聚类微调一个大语言模型智能体,以复现该聚类的标注行为,并通过结合个体和团队层面奖励的偏好优化来协调这些智能体。
We evaluate MAP-PO in four settings defined by two languages and two backbone language models, asking whether each agent reproduces the annotations of its own cluster and whether the agents together reproduce the majority label. Two findings hold in all four settings. First, without fine-tuning the agents behave almost identically, so cluster-specific training is necessary. Second, we show that training each agent only on the labels of its own cluster pushes the agents far beyond the clusters they should represent, while adding a shared team-level training signal consistently keeps each agent calibrated to its cluster.
我们在由两种语言和两种骨干语言模型定义的四种设置中评估了 MAP-PO,旨在考察每个智能体是否能复现其所属聚类的标注,以及这些智能体共同是否能复现多数投票标签。在所有四种设置中,我们得出了两个结论。首先,未经微调的智能体表现几乎完全相同,因此针对特定聚类的训练是必要的。其次,我们发现仅在各自聚类的标签上训练智能体会导致其偏离应代表的聚类,而增加共享的团队级训练信号则能持续使每个智能体保持在其聚类的校准范围内。
Paper Details:
- Authors: Hadi Mohammadi, Tina Shahedi, Robert A. Bagheri, Mehdi Dastani, Masoume M. Raeissi
- Subjects: Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (cs.LG)
- arXiv ID: 2608.04056
论文详情:
- 作者: Hadi Mohammadi, Tina Shahedi, Robert A. Bagheri, Mehdi Dastani, Masoume M. Raeissi
- 学科: 计算与语言 (cs.CL);计算机与社会 (cs.CY);机器学习 (cs.LG)
- arXiv ID: 2608.04056