Not All Memories Are Equal: Hierarchical Collaborative Memory for Validity-Aware Retrieval in LLM Agents
Not All Memories Are Equal: Hierarchical Collaborative Memory for Validity-Aware Retrieval in LLM Agents
并非所有记忆都生而平等:用于 LLM 智能体有效性感知检索的分层协作记忆
Abstract: In team collaboration scenarios, memory is heterogeneous and continually evolving. Team memories capture collective decisions, protocols, and current consensus, while individual memories preserve member-specific observations, execution traces, and intermediate progress. 摘要: 在团队协作场景中,记忆具有异构性且在不断演变。团队记忆捕捉集体决策、协议和当前共识,而个人记忆则保留了成员特定的观察结果、执行轨迹和中间进度。
Existing memory-augmented systems typically retrieve from all stored memories as a flat pool, ranking them by semantic relevance, importance, or recency without modeling hierarchical structure or evolving validity. As a result, they often surface semantically relevant but outdated or conflicting memories, especially individual memories that no longer align with current team consensus, instead of prioritizing currently valid memories. 现有的记忆增强系统通常将所有存储的记忆视为一个扁平的池进行检索,通过语义相关性、重要性或时效性进行排序,而没有对分层结构或演变的有效性进行建模。因此,它们往往会提取出语义相关但已过时或冲突的记忆(尤其是那些不再符合当前团队共识的个人记忆),而不是优先考虑当前有效的记忆。
This is particularly problematic when collaborative LLM agents answer user questions, since their responses should be grounded in valid memories. We propose HiCoMER, a framework for hierarchical collaborative memory management and validity-aware retrieval in LLM agents. 当协作式 LLM 智能体回答用户问题时,这尤其成问题,因为它们的回答应基于有效的记忆。我们提出了 HiCoMER,这是一个用于 LLM 智能体中分层协作记忆管理和有效性感知检索的框架。
HiCoMER first maintains the validity of team and individual memories and then retrieves memories that remain valid, rather than retrieving directly from all stored memories. It consists of three components: a Hierarchical Memory Conflict Updater, a Validity-Aware Memory Retriever, and a Memory-Grounded Answer Generator. HiCoMER 首先维护团队和个人记忆的有效性,然后仅检索仍然有效的记忆,而不是直接从所有存储的记忆中进行检索。它由三个组件组成:分层记忆冲突更新器、有效性感知记忆检索器和基于记忆的答案生成器。
To evaluate HiCoMER, we construct two new datasets for memory-grounded question answering in collaborative settings. Experiments on both datasets show that HiCoMER consistently outperforms strong baselines by reducing outdated retrieval, preserving current team consensus, and improving downstream QA quality. 为了评估 HiCoMER,我们构建了两个用于协作环境下基于记忆的问答的新数据集。在两个数据集上的实验表明,HiCoMER 通过减少过时信息的检索、保持当前团队共识并提高下游问答质量,始终优于强基准模型。