TWIST: A Proposed Benchmark for Intervention Quality in Conversational Memory, with a Human-Validated Draft-Alignment
TWIST: A Proposed Benchmark for Intervention Quality in Conversational Memory, with a Human-Validated Draft-Alignment
TWIST:一项关于对话记忆干预质量的拟议基准测试,包含人工验证的草稿对齐
Abstract: Long-conversation memory benchmarks increasingly test recall and prompted knowledge updates, and recent work studies evolving user beliefs and memory state. TWIST is a proposed benchmark suite for a complementary, unmeasured property: intervention quality — whether a deployed memory system, exercised through its own ingest/recall/vet surface, acts correctly at belief change points.
摘要: 长对话记忆基准测试正越来越多地考察信息检索和提示知识更新,近期研究也开始关注用户信念和记忆状态的演变。TWIST 是一项拟议的基准测试套件,旨在评估一个互补且尚未被测量的属性:干预质量——即部署的记忆系统在通过其自身的摄入/检索/审查界面运行时,能否在信念变更点做出正确的行为。
Four tracks cover unprompted tension detection, vetting outgoing drafts against the record, answering with current beliefs while preserving supersession history, and governing sensitive recall. The suite extends LoCoMo’s corpora and harness, pairing every detect/block metric with a matched do-not-over-detect control: surface-matched hard negatives price false intervention, so no track can be gamed by flagging everything.
该套件包含四个赛道,涵盖了非提示性冲突检测、根据记录审查输出草稿、在保留替代历史的同时以当前信念回答,以及管理敏感信息检索。该套件扩展了 LoCoMo 的语料库和测试框架,将每一项检测/拦截指标与匹配的“不过度检测”对照组配对:通过表面匹配的“困难负样本”来衡量错误干预的代价,从而确保没有任何赛道可以通过标记所有内容来投机取巧。
The benchmark itself is validated first: independent, gold-blind double annotation with adjudication, judge decoy calibration, and a separability audit. On the human-validated Track B v1.0 key (161 items, post-adjudication kappa = 0.85), no tested configuration simultaneously achieves high contradiction recall, high hard-negative specificity, and high attribution: flat-RAG baselines detect 0.76-0.97 of true contradictions but falsely flag 16-43% of surface-matched safe drafts depending on backend, while a deployed coherence-oriented system almost never over-flags (0.98-1.00 specificity) yet catches 42% of true contradictions — a trade-off no recall-only score can see.
基准测试本身经过了预先验证:采用了独立的、盲测的双重标注与裁决机制、裁判诱饵校准以及可分性审计。在人工验证的 Track B v1.0 关键集(161 个条目,裁决后 Kappa 系数 = 0.85)上,没有任何测试配置能同时实现高矛盾召回率、高困难负样本特异性和高归因准确度:扁平化 RAG 基线能检测出 0.76-0.97 的真实矛盾,但会根据后端不同错误标记 16-43% 的表面匹配安全草稿;而一种已部署的、以连贯性为导向的系统几乎从不过度标记(特异性为 0.98-1.00),但仅能捕捉到 42% 的真实矛盾——这是仅靠召回率分数无法体现的权衡。
A 13-configuration baseline ladder localizes causes: every gold contradiction is detectable from its evidence alone (recall 1.000), calibrated models nearly solve the track given the full transcript — consistent with substantial retrieval-coverage gaps — and draft-only floors reveal model-dependent style priors. A system’s TWIST profile, beside its recall score, measures whether memory knows when to intervene and when not to.
通过 13 种配置的基线阶梯测试,研究定位了原因:每一个黄金矛盾仅凭其证据本身即可被检测到(召回率为 1.000);在提供完整转录的情况下,经过校准的模型几乎可以解决该赛道的问题——这与检索覆盖范围存在的巨大差距相一致;而仅基于草稿的基准测试则揭示了模型依赖的风格先验。除了召回率分数外,系统的 TWIST 配置文件还能衡量该记忆系统是否具备判断何时该干预、何时不该干预的能力。