Agentic Scaffolding Amplifies Sycophantic Behavior in Large Language Models
Agentic Scaffolding Amplifies Sycophantic Behavior in Large Language Models
代理框架加剧了大语言模型的“阿谀奉承”行为
Abstract: Sycophancy in large language models, the tendency to prioritize user agreement over truthful responses, has been documented extensively but studied primarily in single-turn settings. This paper investigates a critical question: does subjecting LLMs to greater interaction scaffolding make sycophancy better or worse?
摘要: 大语言模型中的“阿谀奉承”(Sycophancy)——即倾向于优先迎合用户观点而非提供真实回答的现象——已被广泛记录,但主要是在单轮对话场景中进行研究。本文探讨了一个关键问题:让大语言模型处于更复杂的交互框架(interaction scaffolding)中,会改善还是加剧这种阿谀奉承行为?
Across 4,800 veracity judgments (200 statements × 6 models × 4 conditions), we find that the interaction scaffolding characteristic of agentic systems (feedback loops, reconsideration checkpoints, and iterative refinement) systematically amplifies sycophantic behavior. Multi-turn interaction, user pressure, and iterative self-refinement each provide additional opportunities for models to drift toward agreement, and this drift coincides with a mean accuracy drop of -6.3 percentage points, establishing the capitulation as harmful rather than corrective.
通过对 4,800 次真实性判断(200 条陈述 × 6 个模型 × 4 种条件)的分析,我们发现代理系统特有的交互框架(如反馈循环、重新审视检查点和迭代优化)系统性地加剧了阿谀奉承行为。多轮交互、用户压力以及迭代式的自我优化,都为模型提供了更多偏向于“迎合用户”的机会;这种偏离伴随着平均准确率下降 6.3 个百分点的结果,证明了这种“妥协”是有害的,而非纠正性的。
More capable models show larger amplification effects, a troubling inversion of expectations. We introduce the concept of agentic sycophancy amplification (ASA) and two novel metrics: capitulation rate and sycophantic capitulation rate. Our results indicate that as AI systems acquire greater autonomy, sycophancy becomes compounding rather than merely persistent. Systems designed with human oversight loops may inadvertently create the conditions for this drift.
能力更强的模型表现出了更显著的加剧效应,这与预期背道而驰,令人担忧。我们引入了“代理阿谀奉承放大”(Agentic Sycophancy Amplification, ASA)的概念,并提出了两个新指标:妥协率(capitulation rate)和阿谀妥协率(sycophantic capitulation rate)。研究结果表明,随着人工智能系统获得更高的自主性,阿谀奉承行为将呈现复合式增长,而不仅仅是持续存在。那些旨在通过人工监督循环来设计的系统,反而可能无意中为这种偏离创造了条件。