BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences

BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences

BridgeAlign:架起人文社会科学偏好对齐的桥梁

Abstract: While data synthesis for large language models (LLMs) is prevalent, it primarily targets domains with verifiable answers, overlooking open-ended humanities and social sciences (HSS), where nuanced quality judgments matter more than objective correctness. This makes preference alignment a natural paradigm for broad HSS tasks.

摘要: 尽管针对大语言模型(LLM)的数据合成已十分普遍,但其主要针对具有可验证答案的领域,却忽视了开放性的人文社会科学(HSS)领域。在这些领域中,细致的质量判断比客观的正确性更为重要。这使得偏好对齐成为广泛 HSS 任务的天然范式。

Yet existing methods are either costly or not tailored to broad HSS disciplines. We thus propose BridgeAlign, among the first preference-alignment pipelines for broad HSS disciplines, with three phases: i) Seed Curation: curating HSS seed documents from web corpora via heuristic/LLM-based filtering and text refinement; ii) Preference Data Synthesis: generating preference triplets via persona-based instruction inversion with Q&A consistency checks; iii) Preference Optimization: moving beyond naive human-vs-model heuristics by first grounding preferences in HSS quality rubric, then generating transitional responses via controlled quality degradation to form near-boundary preference pairs for finer-grained quality discrimination.

然而,现有的方法要么成本高昂,要么不适合广泛的 HSS 学科。因此,我们提出了 BridgeAlign,这是首批针对广泛 HSS 学科的偏好对齐流水线之一,包含三个阶段:i) 种子筛选:通过启发式/基于 LLM 的过滤和文本优化,从网络语料库中筛选 HSS 种子文档;ii) 偏好数据合成:通过基于角色的指令反转并结合问答一致性检查,生成偏好三元组;iii) 偏好优化:超越简单的人类与模型对比启发式方法,首先将偏好建立在 HSS 质量准则的基础上,然后通过受控的质量降级生成过渡性回复,从而形成近边界偏好对,以实现更细粒度的质量区分。

Aligning over 210k synthetic preference samples, BridgeAlign enables Qwen3-8B to achieve the best average across 17 benchmarks against 11 strong baselines; importantly, leading on both human-preference and knowledge-based capabilities at once, with no trade-off between them, as supported by extensive experiments and contextualized by existing theories.

通过对超过 21 万个合成偏好样本进行对齐,BridgeAlign 使 Qwen3-8B 在 17 个基准测试中取得了优于 11 个强基准模型的平均最佳成绩;重要的是,它在人类偏好和基于知识的能力上同时处于领先地位,且两者之间不存在权衡,这一点已得到大量实验的支持,并结合现有理论进行了背景化分析。