Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language Models

Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language Models

超越对称智能体:小语言模型中的认知多样性与多智能体辩论

Abstract: Multi-agent debate (MAD) reportedly improves reasoning and factuality over single-model inference, but prior work treats agents as symmetric peers, leaving open what drives the gains. We test the hypothesis that cognitive diversity among agents is the driver, in the setting where the question is still measurable: small open-weight models with benchmark headroom.

摘要: 据报道,多智能体辩论(MAD)在推理和事实准确性方面优于单模型推理,但以往的研究将智能体视为对称的对等体,这使得驱动这些收益的因素尚不明确。我们在问题仍可量化的环境下测试了“智能体间的认知多样性是驱动因素”这一假设,即使用具有基准测试余量(benchmark headroom)的小型开源权重模型进行实验。

Across 23 models from eleven vendor families, five tasks, and 5,500+ debate and control runs, we vary diversity along three axes - personas, sampling temperature, and model identity - pairing every debate configuration with a generation-budget-matched majority-vote control. The hypothesis is rejected on every axis.

通过涵盖来自 11 个厂商系列的 23 个模型、5 项任务以及 5,500 多次辩论和对照运行,我们从角色设定(personas)、采样温度(sampling temperature)和模型身份(model identity)三个维度改变多样性,并将每种辩论配置与生成预算匹配的多数投票(majority-vote)对照组进行配对。实验结果在所有维度上都否定了该假设。

Debate beats single-agent inference (3—7 points where tasks have headroom) but at matched budget conditions it ties or even loses to self-consistency sampling at 1.6$\times$ the wall-clock and 3.4$\times$ the token cost. Persona prompting reduces accuracy and a dose-response experiment over each model’s full combinatorial persona space shows the cost is a persona tax, not a diversity tax: redundant personas hurt most, while maximally-diverse teams recover part of the loss.

辩论确实优于单智能体推理(在任务有余量的情况下提升 3-7 个百分点),但在预算匹配的条件下,它与自洽性采样(self-consistency sampling)持平甚至落后,且耗时是后者的 1.6 倍,Token 成本是后者的 3.4 倍。角色提示(Persona prompting)会降低准确率,针对每个模型完整组合角色空间进行的剂量反应实验表明,这种成本是“角色税”而非“多样性税”:冗余的角色设定损害最大,而最大程度多样化的团队可以挽回部分损失。

Furthermore, mixed-model teams lose to majority votes over their own rosters, with accuracy tracking member capability rather than heterogeneity, and nearly all of debate’s benefit comes from the first exchange of answers. We further identify a pervasive measurement hazard in which debate transcripts silently overflow serving context windows, whose correction alone moves our debate-versus-sampling comparison from $-1.8$ points to parity.

此外,混合模型团队的表现不如其成员各自的多数投票结果,准确率取决于成员的能力而非异质性,且辩论带来的几乎所有收益都来自第一轮答案交换。我们还发现了一个普遍存在的测量隐患:辩论记录会悄无声息地溢出服务上下文窗口,仅修正这一问题就使我们的辩论与采样对比结果从 -1.8 个百分点提升至持平。

Our results recast reported MAD gains as an ensemble-sampling effect and provide the budget-matched, contamination-checked baseline bar that future debate mechanisms should be required to clear.

我们的研究结果将此前报道的 MAD 收益重新定义为一种集成采样效应(ensemble-sampling effect),并提供了未来辩论机制必须达到的、经过预算匹配和污染检查的基准线。