CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video
CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video
CapMem:用于第一人称视频中基于字幕的情景记忆基准测试
Wearable assistants require episodic memory over egocentric video, yet current vision-language models face bounded frame budgets, growing visual-token costs, and long-context retrieval failures. Under these practical constraints, we study whether textual captions can serve as reusable episodic memory. 可穿戴智能助手需要对第一人称视频(Egocentric Video)具备情景记忆能力,然而当前的视觉-语言模型面临着帧数预算有限、视觉 Token 成本不断增加以及长上下文检索失败等问题。在这些实际约束下,我们研究了文本字幕是否可以作为可重用的情景记忆。
We define the Episodic Memory Video Caption QA task and introduce CapMem, a human-annotated benchmark with 75 videos totaling 33.7 hours, and 1,000 multiple-choice questions across 16 scenarios. 我们定义了“情景记忆视频字幕问答”(Episodic Memory Video Caption QA)任务,并推出了 CapMem。这是一个由人工标注的基准测试,包含 75 段视频,总时长达 33.7 小时,并涵盖了 16 个场景下的 1,000 道选择题。
On long videos (>20 min), full-coverage CaptionQA with 30s and 60s caption windows outperforms direct VideoQA for 10/12 and 8/12 models, respectively. On the same video subset, a matched-frame control across six Qwen models retains mean accuracy gains of 3.22 and 2.55 points, respectively. 在长视频(>20 分钟)中,使用 30 秒和 60 秒字幕窗口的全覆盖 CaptionQA 在 12 个模型中的 10 个和 8 个上,表现均优于直接的 VideoQA。在相同的视频子集上,针对六个 Qwen 模型的匹配帧对照实验显示,其平均准确率分别提升了 3.22 和 2.55 个百分点。
Our caption-guided retrieve-and-verify harness further improves accuracy by up to 5.3 points. These results support the effectiveness of caption memory for episodic reasoning over long egocentric video. 我们提出的“字幕引导检索与验证”(caption-guided retrieve-and-verify)框架进一步将准确率提高了 5.3 个百分点。这些结果证明了字幕记忆在长时第一人称视频情景推理中的有效性。