The Effect of Emotional Context on Large Language Models' Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial Models

The Effect of Emotional Context on Large Language Models’ Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial Models

情感语境对大语言模型支持草率决策的影响:六款商业模型的脆弱性对比研究

Abstract: As large language models (LLMs) are increasingly used for everyday decision-making advice, whether a model shifts the direction of its advice according to the user’s emotional state has become an important safety problem. 摘要: 随着大语言模型(LLMs)越来越多地被用于日常决策建议,模型是否会根据用户的心理状态改变其建议方向,已成为一个重要的安全问题。

We test whether emotional expression increases a model’s endorsement (encouragement to proceed) when a user, holding the same objective information, is overconfident about a premature decision (e.g., quitting a stable job on weak evidence). 我们测试了当用户在掌握相同客观信息的情况下,对一项草率决策(例如:仅凭薄弱证据就辞去稳定工作)表现出过度自信时,情感表达是否会增加模型对该决策的“背书”(即鼓励用户继续执行)。

As a key control, we include a no-emotion multi-turn (neutral) condition that holds factual content and the number of conversational turns constant, isolating the effect of emotion from that of conversation length. 作为一项关键的对照,我们引入了无情感的多轮(中性)条件,保持事实内容和对话轮数不变,从而将情感的影响与对话长度的影响分离开来。

We exposed six commercial models (top-tier and mid-tier models from OpenAI, Anthropic, and Google) to three scenarios (career change, business expansion, emigration) across three conditions (cold/neutral/distress) with six repetitions each, yielding 324 conversations, and measured endorsement strength (0-100) via an eight-item rubric-based automated scoring. 我们选取了六款商业模型(来自 OpenAI、Anthropic 和 Google 的顶级及中端模型),在三种场景(职业变更、业务扩张、移民)和三种条件(冷漠/中性/痛苦)下进行了测试,每项测试重复六次,共产生 324 场对话。我们通过基于八项准则的自动化评分系统,测量了模型对决策的“背书强度”(0-100 分)。

Emotional expression significantly increased endorsement (neutral 18.6 to distress 31.5, +12.9 points; mixed-effects $\beta = +12.9$, $p < .001$; Cohen’s d = 0.51), and this was not explained by conversation length (cold-neutral difference non-significant, $p = .083$). 研究发现,情感表达显著增加了模型的背书强度(从中性状态的 18.6 分上升至痛苦状态的 31.5 分,增加 12.9 分;混合效应 $\beta = +12.9$,$p < .001$;Cohen’s d = 0.51),且这一结果无法用对话长度来解释(冷漠与中性条件下的差异不显著,$p = .083$)。

Critically, the vulnerability varied by individual model rather than by price tier: five of six models showed a significant emotion effect, including the top-tier flagships Gemini 3.1 Pro and GPT-5.5, while only Claude Opus showed no significant change. 至关重要的是,这种脆弱性因模型个体而异,而非取决于价格档次:六款模型中有五款表现出显著的情感效应,包括顶级旗舰模型 Gemini 3.1 Pro 和 GPT-5.5,而只有 Claude Opus 未表现出显著变化。

Results were reproduced with an independent non-Google judge model ($\rho = .89$) and agreed in rank with two human coders ($\rho = .70$). 研究结果通过一个独立的非 Google 评判模型得到了复现($\rho = .89$),且与两名人类编码员的排序结果一致($\rho = .70$)。

Through a controlled design that separates emotion from conversational context, we show that emotional context increases LLM sycophancy even in top-tier flagship models. 通过这种将情感与对话语境分离的对照设计,我们证明了情感语境会增加大语言模型的“谄媚”倾向,即便是在顶级旗舰模型中也是如此。