Rater State Bias in RLHF Preference Data: An Audit Framework
Rater State Bias in RLHF Preference Data: An Audit Framework
RLHF 偏好数据中的评估者状态偏差:一种审计框架
Abstract: We identify a structured confound in Reinforcement Learning from Human Feedback (RLHF). Pairwise preference labels are intended to reflect the compared outputs, but they may also reflect the rater’s state during annotation. Under sustained stressful or distressing conditions, raters’ preferences may shift over time. As a result, preference data can encode rater state alongside judgments about response quality.
摘要: 我们在人类反馈强化学习(RLHF)中发现了一种结构性混杂因素。成对偏好标签本意是反映被比较的输出结果,但它们也可能反映了评估者在标注过程中的状态。在持续的压力或痛苦条件下,评估者的偏好可能会随时间发生偏移。因此,偏好数据不仅包含对响应质量的判断,还可能编码了评估者的状态。
These shifts differ from ordinary disagreement or random label noise. They are state dependent, can be shared across annotators working under similar conditions, and can propagate through reward modeling and policy optimization. We therefore propose rater state shift as a plausible and testable source of structured bias in RLHF preference data.
这些偏移不同于普通的意见分歧或随机标签噪声。它们具有状态依赖性,可以在相似工作条件下的评估者之间共享,并可能通过奖励建模和策略优化过程进行传播。因此,我们提出“评估者状态偏移”是 RLHF 偏好数据中一种合理且可测试的结构性偏差来源。
This paper develops a hypothesis and an audit framework for studying this source of bias. We define rater state shift, rater state confound, and correlated rater state bias. We also define survival level emotional authenticity as a measurable response pattern using lexical, pragmatic, discourse, and safety related features.
本文针对这一偏差来源提出了假设并构建了审计框架。我们定义了评估者状态偏移、评估者状态混杂以及相关评估者状态偏差。我们还利用词汇、语用、话语和安全相关特征,将“生存级情感真实性”(survival level emotional authenticity)定义为一种可测量的响应模式。
We analyze how correlated rater state bias can survive aggregation and enter learned reward signals. We derive five falsifiable predictions and effect size thresholds for an initial audit. Finally, we present an audit protocol and pilot study plan that can be applied to publicly available instruction tuned models. We do not infer the training history of any specific deployed model. Our goal is to isolate a plausible and testable source of structured bias in RLHF preference data.
我们分析了相关评估者状态偏差如何能够在聚合过程中留存并进入学习到的奖励信号中。我们推导了五个可证伪的预测及用于初步审计的效应量阈值。最后,我们提出了一套可应用于公开指令微调模型的审计协议和试点研究计划。我们不对任何特定已部署模型的训练历史进行推断。我们的目标是分离出 RLHF 偏好数据中一种合理且可测试的结构性偏差来源。