More Is Not More: What Matters for Diversity in LLM Opinions?

More Is Not More: What Matters for Diversity in LLM Opinions?

越多并非越好:什么才是影响大语言模型观点多样性的关键?

Abstract: Large language models are increasingly used to simulate diverse human opinions in open-ended tasks such as synthetic surveys, focus group modeling, and public opinion prediction. However, LLM outputs exhibit systematic opinion homogenization. Practitioners have explored various interventions to increase diversity, but the landscape remains fragmented: different methods are evaluated in isolation with incomparable metrics, and in practice they are typically deployed and upgraded simultaneously, making it difficult to attribute gains to specific components.

摘要: 大语言模型(LLM)正越来越多地被用于在开放式任务中模拟人类的多样化观点,例如合成调查、焦点小组建模和舆论预测。然而,LLM 的输出表现出系统性的观点同质化。从业者已经探索了各种干预措施来增加多样性,但目前的研究领域仍然碎片化:不同的方法在评估时彼此孤立且指标无法比较;在实践中,这些方法通常被同时部署和升级,导致难以将性能提升归因于特定的组件。

To advance a more scientific understanding of LLM output diversity, we design a factorial experiment that separates two primary intervention dimensions: input conditioning (operationalized through persona depth) and interaction architecture. We evaluate all conditions on 100 real-user open-ended questions across 7 models, measuring diversity with multiple complementary metrics.

为了对 LLM 输出的多样性建立更科学的理解,我们设计了一个析因实验,将两个主要的干预维度分离开来:输入条件设定(通过角色深度实现)和交互架构。我们在 7 个模型上针对 100 个真实用户的开放式问题评估了所有条件,并使用多个互补指标来衡量多样性。

Our findings challenge several common assumptions. First, more persona detail does not monotonically increase diversity. The initial step of persona conditioning already captures the majority of the gain, while further elaboration with demographic detail does not consistently improve and can reduce diversity on some models.

我们的研究结果挑战了几个常见的假设。首先,更多的角色细节并不会单调地增加多样性。角色设定的初始步骤已经捕捉到了大部分的增益,而进一步添加人口统计学细节并不能持续改善多样性,在某些模型上甚至可能降低多样性。

Second, rather than seeking a single best interaction architecture, we find that different architectures explore largely non-overlapping opinion regions. Combining multiple architectures yields broader coverage than optimizing any one. Third, commonly attempted low-cost alternatives such as raising sampling temperature and adding diversity instructions produce negligible effects compared to structured interventions.

其次,我们发现与其寻找单一的最佳交互架构,不如认识到不同的架构探索的是很大程度上互不重叠的观点区域。结合多种架构比优化单一架构能产生更广泛的覆盖范围。第三,与结构化干预措施相比,常见的低成本替代方案(如提高采样温度和添加多样性指令)所产生的影响微乎其微。

Overall, our work demonstrates that diversity is not a product of scaling along any single dimension, but is highly sensitive to the structural form and combination of interventions.

总的来说,我们的研究表明,多样性并非沿单一维度扩展的产物,而是对干预措施的结构形式和组合高度敏感。