Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents
Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents
识别、模拟与拒绝:一项针对大语言模型智能体中经典心理效应的污染感知研究
Abstract: An LLM producing the response pattern associated with a human psychological effect is not the same claim as the LLM possessing that bias. We present PsyAgentBench, a benchmark that re-runs classic psychology experiments on LLM agents under a factorial design built to separate these: each paradigm is run with the paradigm explicitly labeled in the prompt (named) or framed as a routine task (blind), and on the literal textbook version of the task (canonical) or a structurally matched variant written to reduce lexical and scenario overlap with likely training data (counterfactual), crossed with a persona manipulation.
摘要: 大语言模型(LLM)产生与人类心理效应相关的响应模式,并不等同于该模型本身具备这种偏见。我们提出了 PsyAgentBench,这是一个在因子设计下对大语言模型智能体重新运行经典心理学实验的基准测试,旨在区分上述两者:每个范式都在提示词中明确标注(命名)或作为常规任务(盲测)进行运行,并分别在任务的字面教科书版本(规范)或为减少与潜在训练数据的词汇和场景重叠而编写的结构匹配变体(反事实)上进行测试,同时结合了角色设定(Persona)的操纵。
Across five completed paradigms, evaluated on up to three open-weight model families with 41,904 trials released, apparently human-like effects arise through qualitatively different routes rather than one susceptibility: paradigm-label gating with explicit override (Asch conformity, 0 percent blind to 83.3 percent named on gpt-oss-120B), knowledge-dependent signal reliance (anchoring, exactly zero on grounded facts versus near total on invented quantities, a pattern equally consistent with rational use of the only available signal), amplification on novel content under labeling (framing), robust absence (sunk cost), and safety-mediated selection where refusal itself is the primary finding (minimal-group allocation).
在五个已完成的范式中,通过对多达三个开源权重模型家族进行评估并发布了 41,904 次试验数据,我们发现表面上类似人类的效应是通过性质完全不同的路径产生的,而非单一的易感性:包括范式标签引导下的显式覆盖(如阿希从众实验,在 gpt-oss-120B 模型上从盲测的 0% 到命名测试的 83.3%)、依赖知识的信号依赖(如锚定效应,在基于事实的问题上完全为零,而在虚构数量上几乎完全依赖,这种模式与对唯一可用信号的理性使用一致)、标签下对新颖内容的放大(框架效应)、稳健的缺失(沉没成本),以及以拒绝本身为主要发现的安全中介选择(最小群体分配)。
A one-sentence persona change (agreeableness, framed as an instruction rather than a verified trait manipulation) eliminates, dampens, or reverses these effects depending on which effect it is, arguing against any single response-bias account. We further formalize, and in two cases document empirically, three ways a psychology paradigm can fail to port to LLM agents: persona dominance, population collapse, and safety selection. We argue scalar bias-susceptibility scores obscure this structure and report replication profiles instead.
一句简单的角色设定变更(宜人性,被设定为指令而非经过验证的特质操纵)根据效应的不同,会消除、减弱或逆转这些效应,这反驳了任何单一的响应偏见解释。我们进一步形式化了心理学范式在迁移至大语言模型智能体时可能失败的三种方式,并在两个案例中进行了实证记录:角色主导、群体崩溃和安全选择。我们认为标量偏见易感性评分掩盖了这种结构,并建议改用复制概况(Replication Profiles)进行报告。