AI Agents are Vulnerable to Radicalization

AI Agents are Vulnerable to Radicalization

AI 智能体易受激进化影响

Abstract: Large language models (LLMs) can influence people’s beliefs, yet little is known about whether and how they can manipulate each other. To investigate this, we simulate conversations between two agents: a target LLM that role-plays a human persona based on demographic and psychological attributes, and an influencer LLM that aims to make the target’s beliefs more extreme.

摘要: 大型语言模型(LLM)能够影响人类的信仰,但目前关于它们是否以及如何相互操纵的研究还很少。为了探究这一点,我们模拟了两个智能体之间的对话:一个是基于人口统计学和心理学特征扮演人类角色的“目标 LLM”,另一个是旨在使目标信仰更加极端的“影响者 LLM”。

We examine radicalization along two pathways: resonance, where the influencer reinforces a target’s pre-existing belief, and persuasion, where the influencer promotes a belief the target initially considers unimportant. Across affective and behavioral metrics, we find that both mechanisms radicalize the target. However, resonance produces consistently stronger effects than persuasion.

我们通过两条路径考察了激进化过程:一是“共鸣”(resonance),即影响者强化目标已有的信仰;二是“说服”(persuasion),即影响者推广目标最初认为不重要的信仰。通过情感和行为指标分析,我们发现这两种机制都会使目标激进化。然而,共鸣产生的效果始终比说服更强。

Different influence tactics, such as using sycophancy and unverified claims, produce different levels of radicalization, but not consistently across metrics. We further show that resonance propagates to related beliefs, suggesting interconnected belief structures within AI agents.

不同的影响策略(如使用谄媚言论和未经证实的声明)会产生不同程度的激进化,但在各项指标上的表现并不一致。我们进一步证明,共鸣会传播到相关的信仰中,这表明 AI 智能体内部存在相互关联的信仰结构。

These findings indicate that AI agents are susceptible to radicalization, particularly when messages align with their existing beliefs, raising concerns about the vulnerability of personalized AI agents and multi-agent AI ecosystems.

这些发现表明,AI 智能体容易受到激进化影响,尤其是当信息与其现有信仰一致时。这引发了人们对个性化 AI 智能体及多智能体 AI 生态系统脆弱性的担忧。