Faking Good and Faking Bad in LLMs: Response Distortion Across Dark Triad Personality Traits

Faking Good and Faking Bad in LLMs: Response Distortion Across Dark Triad Personality Traits

大语言模型中的“伪装好”与“伪装坏”:黑暗三角人格特质的响应偏差研究

Abstract: Social desirability and impression management are pervasive sources of response distortion in human personality assessment, yet their effects on Large Language Models (LLMs) remain underexplored. This study investigates whether contemporary LLMs systematically modulate the expression of Dark Triad traits (Machiavellianism, narcissism, and psychopathy) under fake-good and fake-bad conditions.

摘要: 社会赞许性和印象管理是人类人格评估中普遍存在的响应偏差来源,但它们对大语言模型(LLM)的影响仍未得到充分研究。本研究探讨了当代大语言模型在“伪装好”(fake-good)和“伪装坏”(fake-bad)条件下,是否会系统性地调节其黑暗三角人格特质(马基雅维利主义、自恋和精神病态)的表现。

Seven state-of-the-art models were evaluated across two ecologically relevant contexts: employment selection and forensic evaluation, in which socially desirable or undesirable incentives were conveyed through contextual framing. Trait expression was measured using standard psychometric scoring procedures and compared with self-assessment baselines at both aggregate and item levels.

研究评估了七种最先进的模型,涵盖了两个具有生态相关性的场景:就业筛选和司法评估。在这些场景中,通过情境框架传达了社会赞许或不赞许的激励因素。研究使用标准的心理测量评分程序来衡量特质表现,并将其与聚合层面和项目层面的自我评估基准进行了比较。

Results revealed systematic and condition-consistent response modulation. Most models reduced Dark Triad scores under fake-good conditions and increased them under fake-bad conditions, although the magnitude and consistency of these effects varied across traits and models. Machiavellianism and narcissism showed the strongest and most coherent shifts, whereas psychopathy displayed greater heterogeneity.

结果显示,模型存在系统性且与条件一致的响应调节。大多数模型在“伪装好”条件下降低了黑暗三角得分,而在“伪装坏”条件下提高了得分,尽管这些效应的幅度和一致性在不同特质和模型之间存在差异。马基雅维利主义和自恋表现出最强且最连贯的偏移,而精神病态则表现出更大的异质性。

Context also influenced responses, with employment scenarios generally producing larger effects than forensic scenarios. An additional experiment showed that explicit fake-bad instructions generated substantially stronger distortions than contextual framing alone.

情境也影响了响应,就业场景通常比司法场景产生更大的效应。一项额外的实验表明,明确的“伪装坏”指令比单纯的情境框架产生了明显更强的偏差。

The results suggest that personality-related outputs should be interpreted in light of the motivational and situational context in which they are elicited. More broadly, they highlight the value of psychometric paradigms for evaluating susceptibility to response distortion, impression management, and context-dependent behavioral shifts, with important implications for LLM benchmarking, alignment evaluation, and robustness assessment.

研究结果表明,在解读与人格相关的输出时,应考虑其产生的动机和情境背景。从更广泛的角度来看,这些发现凸显了心理测量范式在评估响应偏差、印象管理和依赖于情境的行为偏移方面的价值,这对大语言模型的基准测试、对齐评估和鲁棒性评估具有重要意义。