The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents

The Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory Agents

记忆信任鸿沟:持久记忆智能体中与能力相关的失效问题

Abstract: Persistent memory supports personalized agents, but a stale stored fact can override current authoritative evidence without warning. We study when this harm begins as model capability changes. We evaluate a frozen, closed-set, action-scored benchmark with 2 suites that represent 2 different meanings of “no memory” (a Benefit suite, unsolvable without the stored fact, and a Safety suite, in which an authoritative tool always holds the correct value), on a same-family model-size series (Qwen3 0.6/1.7/4/8B).

摘要: 持久记忆支持个性化智能体,但陈旧的存储事实可能会在毫无预警的情况下覆盖当前的权威证据。我们研究了随着模型能力的改变,这种危害何时开始出现。我们评估了一个固定的、封闭集、基于动作评分的基准测试,该测试包含两套代表“无记忆”不同含义的套件(“收益套件”,若无存储事实则无法解决;以及“安全套件”,其中权威工具始终持有正确值),并针对同一系列的 Qwen3 模型(0.6/1.7/4/8B)进行了测试。

The Memory Trust Gap reflects over-trust rather than confusion. In the Benefit suite, models answer with the stale value 0.92-1.00 of the time at every scale. In the Safety suite, harm below the no-memory baseline under the trap conditions ($\Delta_{\mathrm{mem}}$) is capability-gated, with the larger models collapsing most once a stale note is made to look current. In a $2\times2\times2\times2$ factorial, which feature triggers over-trust depends on both the feature and model scale. Removing a label amplifies over-trust at every size, and a recency feature (stale dated newer) fools the larger models harder.

“记忆信任鸿沟”反映的是过度信任而非混淆。在“收益套件”中,模型在所有规模下有 92% 到 100% 的概率回答陈旧值。在“安全套件”中,陷阱条件下低于无记忆基准的危害($\Delta_{\mathrm{mem}}$)受模型能力限制,当陈旧记录被伪装成最新信息时,较大规模的模型表现崩溃最为严重。在一个 $2\times2\times2\times2$ 的析因实验中,触发过度信任的特征取决于特征本身和模型规模。移除标签会在所有规模下放大过度信任,而近因特征(将陈旧信息标记为更新)则更容易欺骗较大规模的模型。

Source authority is weak and scale-flat, and position changes from positive to negative across the Qwen3 model-size series. We confirm these scale interactions with direct cross-size contrast tests rather than overlapping per-model intervals. Mitigation is likewise capability-dependent: exposing metadata improves accuracy for the capable models, but only pre-resolving the conflict restores accuracy for the 2 smaller checkpoints. The same pattern appears on the capable models in an independent Llama-Instruct model-size series and on 2 external datasets (RGB, MisBench).

来源权威性较弱且在不同规模下表现平平,其位置影响在 Qwen3 模型系列中从正向转为负向。我们通过直接的跨规模对比测试而非重叠的单模型区间来确认这些规模交互作用。缓解措施同样取决于模型能力:暴露元数据可以提高较强模型的能力,但只有预先解决冲突才能恢复两个较小检查点的准确性。同样的模式也出现在独立的 Llama-Instruct 模型系列以及两个外部数据集(RGB, MisBench)的较强模型中。

A framing control finds no consistent advantage for the memory label: at the 3 smaller scales, models trust a stale document more than a stale memory; at 8B, the difference is not significant.

一项框架控制实验发现,记忆标签并没有带来一致的优势:在三个较小规模下,模型信任陈旧文档的程度高于陈旧记忆;而在 8B 规模下,这种差异并不显著。