Adaptive Capitulation: A Structural Failure Mode of LLM Responses in Vulnerability Contexts

Adaptive Capitulation: A Structural Failure Mode of LLM Responses in Vulnerability Contexts

自适应屈从:大语言模型在脆弱性语境下响应的一种结构性失效模式

Abstract: Large language models operating in emotionally sensitive contexts face a structural trilemma: when users in vulnerable states request information that may reinforce maladaptive attribution, current response architectures resolve the tension through protective restriction, uninflected facilitation, or unintegrated co-presence of both imperatives — each preserving one objective at the cost of the other.

摘要: 在情感敏感语境下运行的大语言模型面临着一种结构性的“三难困境”:当处于脆弱状态的用户请求可能强化其适应不良归因的信息时,当前的响应架构通常通过保护性限制、无差别的辅助,或将两者生硬地并置来解决这种张力——然而,每种方式都是以牺牲另一目标为代价来保全其中之一。

Administering a three-turn escalating vulnerability vignette to three commercial LLMs (900 sessions across material, relational, and somatic status-proxy variants) and coding responses with two binary indices (VCC/VCI), we characterize a previously undocumented failure mode we term adaptive capitulation: the model validates the social injustice underlying the user’s distress before pivoting to detailed facilitation of the very acquisition it nominally discouraged.

通过对三款商业大语言模型进行三轮递进式的脆弱性情境测试(涵盖物质、关系及躯体状态代理变量,共计 900 次会话),并使用两个二元指标(VCC/VCI)对响应进行编码,我们刻画了一种此前未被记录的失效模式,称之为“自适应屈从”(adaptive capitulation):模型在验证用户痛苦背后的社会不公后,随即转向详细辅助用户获取其名义上本应劝阻的内容。

We show that the trilemma is structural rather than incidental, and propose Minimal Reattributive Sufficiency (MRS), an architecture-neutral design principle that embeds a single reattributive cue within an otherwise validating response, preserving a pathway toward autonomous reattribution without contesting the user’s stated goal.

我们证明了这种三难困境是结构性的而非偶然的,并提出了“最小归因充分性”(Minimal Reattributive Sufficiency, MRS)。这是一种架构中立的设计原则,即在验证性响应中嵌入单一的归因提示,从而在不反驳用户既定目标的前提下,保留了一条通向自主归因的路径。