Which Objectives Need a Dial? Predicting Objective Conflict and Covering Trade-offs in Steerable Pluralistic Alignment

Which Objectives Need a Dial? Predicting Objective Conflict and Covering Trade-offs in Steerable Pluralistic Alignment

哪些目标需要“调节旋钮”?在可控多元对齐中预测目标冲突并覆盖权衡

People hold diverse, sometimes conflicting values, so no single aligned model can satisfy everyone. Pluralistic alignment therefore calls for steerable models that can balance competing objectives differently. Multi-Objective Direct Preference Optimization (MODPO) does this by using an objective weight to span a continuum of trade-offs.

人们持有多样化且有时相互冲突的价值观,因此没有任何单一的对齐模型能够满足所有人。因此,多元对齐(Pluralistic alignment)要求模型具备可控性,能够以不同方式平衡相互竞争的目标。多目标直接偏好优化(MODPO)通过使用目标权重来跨越一系列权衡取舍,从而实现了这一点。

We study two questions: when can one model improve two objectives simultaneously, and how can many trade-offs be covered without training a separate model for each? Across seven objective pairs from HelpSteer and UltraFeedback, two pre-training measurements predict whether objectives align or conflict for human-annotated data, but not for AI-annotated data, where response length and repetition confound reward-model scores.

我们研究了两个问题:一个模型何时能同时改进两个目标?以及如何在不为每个权衡方案单独训练模型的情况下覆盖多种权衡?通过对 HelpSteer 和 UltraFeedback 中的七对目标进行研究,我们发现两项预训练指标可以预测人类标注数据中的目标是一致还是冲突,但对于 AI 标注数据则无效,因为在 AI 数据中,响应长度和重复性会干扰奖励模型的评分。

For broader trade-off coverage, selecting the nearest trained model and merging model parameters both help, but neither consistently matches direct training. These findings yield practical guidance for building steerable models that serve diverse preferences.

为了实现更广泛的权衡覆盖,选择最接近的已训练模型和合并模型参数都有所帮助,但两者都无法始终达到直接训练的效果。这些发现为构建能够服务于多样化偏好的可控模型提供了实践指导。