Gradient-Aligned Pair Selection for Personalized Preference Optimization
Gradient-Aligned Pair Selection for Personalized Preference Optimization
用于个性化偏好优化的梯度对齐配对选择
Personalizing large language models (LLMs) requires aligning generation behavior with user-specific preferences rather than aggregate quality. While Direct Preference Optimization (DPO) provides a stable framework for preference learning, its effectiveness in personalized settings critically depends on how preference pairs are selected. 个性化大语言模型(LLMs)需要将生成行为与用户特定的偏好对齐,而非仅仅追求整体质量。虽然直接偏好优化(DPO)为偏好学习提供了一个稳定的框架,但其在个性化场景下的有效性在很大程度上取决于偏好对的选择方式。
Existing approaches typically rely on heuristic criteria, such as likelihood-based extremes, which decouple optimization from explicit user utility and can lead to degraded personalization. We formalize personalized preference learning as a geometry-aligned optimization problem by analyzing the first-order interaction between gradients of expected user utility and DPO update directions. 现有的方法通常依赖于启发式准则(如基于似然的极值),这使得优化过程与显式的用户效用脱节,并可能导致个性化效果下降。我们通过分析预期用户效用的梯度与 DPO 更新方向之间的一阶交互,将个性化偏好学习形式化为一个几何对齐的优化问题。
Our analysis reveals that, under off-policy sampling, the DPO update transitions from a purely error-corrective signal to a reinforcement-like update when preference margins are directionally aligned with utility gradients. This perspective exposes pair selection as a geometric decision that governs whether preference optimization advances or hinders personalization. 我们的分析表明,在离策略采样(off-policy sampling)下,当偏好边界与效用梯度在方向上对齐时,DPO 更新会从纯粹的纠错信号转变为类似强化学习的更新。这一视角揭示了配对选择是一个几何决策,它决定了偏好优化是促进还是阻碍了个性化。
Motivated by this insight, we propose GAP-DPO (Geometry-Aligned Preference DPO), an iterative algorithm that performs utility-aware, geometry-aligned pair selection while controlling distribution shift via epoch-wise regeneration. Experiments on personalized text generation benchmarks show that GAP-DPO consistently improves stylistic fidelity, preference alignment, and generation quality compared to standard DPO variants. 受此启发,我们提出了 GAP-DPO(几何对齐偏好 DPO),这是一种迭代算法,它在执行效用感知和几何对齐的配对选择的同时,通过周期性的重新生成来控制分布偏移。在个性化文本生成基准测试上的实验表明,与标准的 DPO 变体相比,GAP-DPO 在风格保真度、偏好对齐和生成质量方面均有持续提升。
Together, our results establish gradient alignment as a unifying principle for personalized preference optimization and demonstrate that pair selection is an intrinsic component of the optimization geometry rather than a heuristic preprocessing step. 总之,我们的研究结果确立了梯度对齐作为个性化偏好优化的统一原则,并证明了配对选择是优化几何结构的一个内在组成部分,而非一种启发式的预处理步骤。