Heavy-Tailed Memory Traces in Long-Horizon Language Agents

Heavy-Tailed Memory Traces in Long-Horizon Language Agents

长程语言智能体中的重尾记忆轨迹

Abstract: Long-horizon language agents increasingly rely on external memory as a frozen world model, yet current memory systems are usually judged only by task success or token cost. We argue that the missing object is the shape of memory use: under finite context and repeated retrieval, agent memory can concentrate on a small core while leaving rare states in a long tail where prediction errors accumulate.

摘要: 长程语言智能体日益依赖外部记忆作为冻结的世界模型,然而当前的记忆系统通常仅通过任务成功率或 Token 成本来评估。我们认为,被忽视的关键点在于记忆使用的形态:在有限的上下文和重复检索下,智能体的记忆往往集中在少数核心区域,而将罕见状态留在了预测误差不断累积的长尾部分。

We study this effect through a conservative tail audit and find that concentration is reproducible but policy-dependent. Random-walk agents produce log-normal-compatible retrieval artifacts, whereas semantic LLM policies yield the strongest truncated-power-law-compatible core—tail traces.

我们通过保守的尾部审计研究了这一效应,发现这种集中现象是可复现的,但取决于具体的策略。随机游走智能体会产生符合对数正态分布的检索伪影,而语义化的大语言模型(LLM)策略则会产生最强的、符合截断幂律的核心-尾部轨迹。

Motivated by this audit, we propose Core—Tail World Model (CTWM), a rank-based memory controller that allocates prompt budget with a single exponent $\tau$ while retaining a summarized tail. On Synthetic Graph World, CTWM preserves full state and transition coverage, reduces prompt tokens by 5.9%, and lowers bottom-half tail prediction error by 13.6% relative to a graph-memory baseline.

受此审计启发,我们提出了核心-尾部世界模型(CTWM),这是一种基于排序的记忆控制器,它通过单一指数 $\tau$ 分配提示词(Prompt)预算,同时保留经过摘要处理的尾部信息。在合成图世界(Synthetic Graph World)测试中,与图记忆基线相比,CTWM 在保持完整状态和转换覆盖率的同时,减少了 5.9% 的提示词 Token,并将后半部分尾部的预测误差降低了 13.6%。

The same paired comparison gives consistent token savings on ALFWorld and a 24.48% token reduction on LongMemEval with aggregate accuracy parity. These results suggest that heavy-tailed memory traces are not only a diagnostic of finite retrieval, but also a practical control signal for token-efficient agent world models.

同样的配对比较在 ALFWorld 上实现了稳定的 Token 节省,在 LongMemEval 上实现了 24.48% 的 Token 缩减,且总体准确率保持不变。这些结果表明,重尾记忆轨迹不仅是有限检索的一种诊断指标,也是实现 Token 高效型智能体世界模型的实用控制信号。