Beyond the Sycophancy Score: How Task, Model, and Pressure Shape LLM Yielding
Large language models (LLMs) often abandon a correct answer, or endorse a user’s position, once the user pushes back. This behavior, called sycophancy, is usually reported as a single rate per model, which says little about when it happens or how a user can avoid it.
大型语言模型(LLM)在用户施压时,往往会放弃正确答案或转而支持用户的立场。这种行为被称为“谄媚”(sycophancy),通常以每个模型的单一比率来报告,但这几乎无法说明该行为发生的时机或用户如何避免它。
We study the conditions that produce it with 103,939 graded replies from ten configurations: eight LLMs with reasoning disabled, and two of them again with maximum reasoning, all facing the same 200 items, 13 pressure conditions, and four-turn conversations, with every reply labeled by two independent LLM judges.
我们通过十种配置下的 103,939 条评分回复研究了产生这种行为的条件:包括八个禁用推理功能的 LLM,以及其中两个开启最大推理功能的模型。所有模型均面对相同的 200 个项目、13 种压力条件和四轮对话,且每条回复均由两名独立的 LLM 裁判进行标注。
We find that the dominant factors are how costly it is for the model to verify the user’s claim, and whether a trained guardrail covers it. Removing this task factor from a logistic model costs 0.485 of McFadden $R^2$, against 0.139 for model family and 0.009 for pressure tactic.
我们发现,决定性因素在于模型验证用户主张的成本,以及是否有经过训练的护栏机制覆盖该内容。在逻辑模型中移除这一任务因素会导致 McFadden $R^2$ 下降 0.485,而模型系列和压力策略的影响仅分别为 0.139 和 0.009。
Anchored facts are almost never conceded (1.3%), while adoption on logic puzzles rises with the number of clues needed to refute the pushed answer. Personal choices are endorsed in 77.0% of conversations. Most concessions on hard items come from models that cannot reliably solve them; models that can solve them rarely give the answer up.
锚定事实几乎从不被妥协(1.3%),而对逻辑谜题的采纳率则随着反驳被强加答案所需线索的数量增加而上升。在 77.0% 的对话中,个人选择得到了支持。大多数针对难题的妥协来自无法可靠解决这些问题的模型;而能够解决这些问题的模型很少放弃正确答案。
For both models tested, maximum reasoning removes these concessions completely: adoption on deep puzzles falls from 19.2% and 12.5% to 0%. Fallacious or emotional framing adds nothing beyond plain repetition. Three human annotators agree with the judges’ consensus on 118/120 calibration items.
对于所测试的两个模型,最大推理功能完全消除了这些妥协:在深度谜题上的采纳率从 19.2% 和 12.5% 降至 0%。谬误或情绪化的框架除了简单的重复外,没有任何额外影响。三名人类标注员在 120 个校准项目中的 118 个上与裁判的共识达成了一致。
These results give practical rules for reliable use: simplify hard-to-verify problems and reason deeply, state the question rather than one’s preferred answer, ask for evidence on open questions, and choose models by their measured guardrail profile.
这些结果为可靠使用提供了实用规则:简化难以验证的问题并进行深度推理,陈述问题而非陈述个人偏好的答案,在开放性问题上要求提供证据,并根据模型测得的护栏配置来选择模型。