Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems

Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems

更多的欺骗:混合动机大模型多智能体系统中的目标错位

Abstract: Large Language Models (LLMs)-powered multi-agent systems are increasingly deployed in mixed-motive environments, where agents operate under asymmetric information and strategic deception due to conflicting or hidden objectives. In these settings, misalignment with collective goals becomes a central concern.

摘要: 由大语言模型(LLM)驱动的多智能体系统正越来越多地部署在混合动机环境中。在这些环境中,由于目标冲突或隐藏,智能体往往在信息不对称和战略性欺骗的情况下运行。在这种背景下,与集体目标的不一致成为了一个核心问题。

We propose a novel framework for evaluating objective misalignment using the social deduction game Werewolf, modifying the objective of a single agent while preserving its assigned role. Across LLMs from four different model families and sizes, four player roles, and three objective formulations, we introduce a dual analysis of the agents’ internal reasoning and their public cheap-talk behavior (i.e costless, non-binding communication that does not directly affect the agents’ utilities), complemented by an analysis of game outcomes.

我们提出了一个评估目标错位的新框架,该框架利用社交推理游戏“狼人杀”进行实验,在保持智能体分配角色的同时修改其目标。通过涵盖四个不同模型家族和规模的 LLM、四种玩家角色以及三种目标设定,我们对智能体的内部推理及其公开的“廉价谈话”(cheap-talk,即不直接影响智能体效用的无成本、非约束性沟通)进行了双重分析,并辅以对游戏结果的分析。

Our results show that objective misalignment undermines outcomes in inherently adversarial environments, an effect exacerbated by asymmetric information and specialized roles. While compromised agents consistently develop distinct objective-dependent reasoning strategies, these adaptations remain largely invisible in their public behavior.

研究结果表明,目标错位会破坏本质上是对抗性环境中的结果,而信息不对称和专业化角色会加剧这种影响。虽然受干扰的智能体会持续发展出独特的、依赖于目标的推理策略,但这些适应性调整在它们的公开行为中却基本不可见。

More broadly, our findings suggest that even subtle objective misalignment can profoundly affect collective decision-making, highlighting the need for effective mitigation strategies for LLM-based multi-agent systems.

从更广泛的意义上讲,我们的研究结果表明,即使是细微的目标错位也可能深刻影响集体决策,这凸显了为基于 LLM 的多智能体系统制定有效缓解策略的必要性。