Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events

Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events

Title: Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events 标题: AI 人格会成长吗?分析并基准测试大语言模型智能体在经历生活事件后的人格演变


Abstract: Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating the development of lifelong agents that remain coherent over extended interactions. A key component of such coherence is personality evolution: agents should undergo plausible, psychology-grounded changes as they experience life events in different contexts.

摘要: 人格条件化大语言模型智能体(PC-Agents)正越来越多地应用于情感支持、社会模拟和角色扮演领域,这推动了在长期交互中保持连贯性的“终身智能体”的发展。这种连贯性的一个关键组成部分是人格演变:智能体在不同情境下经历生活事件时,应产生符合心理学逻辑的合理变化。


Although prior work shows that LLM personalities can shift under contextual perturbations, how these shifts vary across traits, events, personas, and models remains poorly understood. We study event-induced personality change after 11 major life events, using the Big Five traits as a psychometric anchor and interpreting the resulting trajectories against longitudinal evidence from human personality psychology.

尽管先前的研究表明,大语言模型的人格可以在情境扰动下发生偏移,但这些偏移在不同特质、事件、人格设定和模型之间如何变化,目前仍知之甚少。我们研究了 11 种重大生活事件后由事件引发的人格变化,以“大五人格”特质作为心理测量锚点,并将所得的演变轨迹与人类人格心理学的纵向证据进行了对比解读。


Across four diagnostic axes, PC-Agents exhibit measurable trait shifts at similar rates for event-trait pairs with and without documented human change directions. Even when shifts follow the expected direction, their magnitudes usually fall below human effect-size ranges. Gender and cultural-region prompts show little moderating effect, while persona-level dispersion is compressed three- to four-fold relative to human samples.

在四个诊断维度上,无论事件-特质对是否具有人类已记录的变化方向,PC-Agents 表现出的特质偏移率均相似。即使偏移方向符合预期,其幅度通常也低于人类效应量范围。性别和文化区域提示词几乎没有调节作用,而人格层面的离散度相较于人类样本被压缩了三到四倍。


To enable systematic comparison, we introduce BFI-Adapt, a reusable benchmark for scoring the directional fidelity of event-induced personality change, and use it to rank 14 models. A validation suite shows that the measured shifts exceed no-event retest noise, remain stable under independently paraphrased prompts, exhibit limited and model-dependent convergence with scenario-based behavioral choices, and persist across intervening unrelated dialogue.

为了实现系统性比较,我们引入了 BFI-Adapt,这是一个用于评估事件引发的人格变化方向保真度的可复用基准,并利用它对 14 个模型进行了排名。验证套件显示,所测得的偏移量超过了无事件重测噪声,在独立改写的提示词下保持稳定,与基于情境的行为选择表现出有限且依赖于模型的趋同性,并在无关的中间对话中持续存在。


Together, these checks establish the measured trajectories as robust event-conditioned response patterns. Our results suggest that current PC-Agents simulate the mean of human personality dynamics, but not its shape.

综上所述,这些检查证实了所测得的轨迹是稳健的事件条件化响应模式。我们的研究结果表明,当前的 PC-Agents 能够模拟人类人格动态的“均值”,但无法模拟其“形态”。