FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents

FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents

FinPerMA:一种基于理论与事件驱动的 LLM 智能体个性化记忆基准测试

Abstract: Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long horizons.

摘要: 大语言模型(LLM)智能体正越来越多地被用作金融咨询等高风险领域的个性化助手,但目前尚不清楚它们是否能够在长期范围内维护和更新个性化的用户模型。

Existing personalized-memory benchmarks primarily test factual retention or rely on weakly constrained model-generated trajectories, leaving event-driven preference adaptation underexplored.

现有的个性化记忆基准测试主要侧重于测试事实留存能力,或依赖于约束较弱的模型生成轨迹,导致对事件驱动的偏好适应性研究不足。

We introduce FinPerMA, an event-grounded benchmark that evaluates personalized memory against frozen longitudinal investor trajectories. Its generation pipeline combines deterministic, theory-informed impact rules, controlled LLM narration, and automated quality screening; a Post-Shock checkpoint isolates whether an agent has integrated a material event into its persistent user model.

我们推出了 FinPerMA,这是一个基于事件的基准测试,旨在通过固定的纵向投资者轨迹来评估个性化记忆。其生成流程结合了确定性的理论影响规则、受控的 LLM 叙述以及自动质量筛选;“冲击后”(Post-Shock)检查点能够隔离并评估智能体是否已将重大事件整合到其持久的用户模型中。

On 2,994 questions from 276 personas, seven frontier LLMs and up to seven memory configurations remain far from saturated: no full-context configuration exceeds approximately 0.47 overall accuracy or approximately 39% on multiple-choice questions.

在针对 276 个人物画像的 2,994 个问题的测试中,七个前沿 LLM 模型及多达七种记忆配置的表现远未达到饱和:没有任何全上下文配置的总体准确率超过约 0.47,或在多项选择题中超过约 39%。

Attribution analysis shows that summary-based memory often preserves factual details while losing the preference signals needed for personalization; simple retrieval can therefore outperform purpose-built memory systems, with the gap widening after shocks.

归因分析表明,基于摘要的记忆往往能保留事实细节,却丢失了个性化所需的偏好信号;因此,简单的检索方法有时反而能优于专门构建的记忆系统,且这种差距在发生重大冲击后会进一步拉大。