The Wisdom of Artificial Deliberative Crowds

The Wisdom of Artificial Deliberative Crowds

人工审议群体的智慧

Abstract: The aggregation of many lay estimates often outperforms individual expert judgment, a phenomenon known as the wisdom of crowds. While this is usually attributed to the independence of estimates, an even stronger effect arises through deliberation: averaging the consensus estimates of small deliberating groups outperforms the classical wisdom of crowds, with individual judgments themselves also becoming more accurate after deliberation.

摘要: 汇总许多普通人的估计往往优于个人的专家判断,这种现象被称为“群体智慧”。虽然这通常归因于估计的独立性,但通过审议(deliberation)会产生更强的效果:对小型审议小组的共识估计进行平均,其表现优于传统的群体智慧,且个人判断在审议后也会变得更加准确。

Whether these improvements transfer to large language models deliberating amongst themselves is unknown. Here we adapt a three-stage deliberation paradigm previously used with human participants for use with large language models from three different families, and test it across four domains of increasing real-world stakes: visual numerical estimation (Study 1), peer review of machine-learning papers (Study 2), detection of hidden malicious behavior by an artificial intelligence agent (Study 3), and sports forecasting against a real prediction market (Study 4).

这些改进是否能迁移到进行相互审议的大型语言模型中,目前尚不清楚。在此,我们将此前用于人类参与者的三阶段审议范式应用于来自三个不同家族的大型语言模型,并在四个现实风险递增的领域进行了测试:视觉数值估计(研究 1)、机器学习论文的同行评审(研究 2)、检测人工智能代理隐藏的恶意行为(研究 3),以及针对真实预测市场的体育预测(研究 4)。

Across domains, deliberation reduced collective error beyond passive aggregation of independent responses, and post-deliberation individual judgments retained this collective gain. Notably, the advantage required model diversity: groups composed of clones of a single model did not benefit from deliberating. These results establish machine deliberation as a general-purpose aggregation mechanism, and point to diversity as an active ingredient.

在各个领域中,审议所减少的集体误差均超过了对独立响应的被动汇总,且审议后的个人判断保留了这种集体收益。值得注意的是,这种优势需要模型的多样性:由单一模型的克隆体组成的群体并不能从审议中获益。这些结果确立了机器审议作为一种通用汇总机制的地位,并指出多样性是其中的关键要素。