Relation Before Entity: Deferred Commitment in Language Model Factual Recall
Relation Before Entity: Deferred Commitment in Language Model Factual Recall
关系先于实体:语言模型事实回忆中的延迟承诺
Abstract: We ask whether relation-type information (e.g., capital-of) and entity-specific information (e.g., France to Paris) become causally active at the final-token position at the same depth during recall.
摘要: 我们探讨了在事实回忆过程中,关系类型信息(例如“……的首都”)和实体特定信息(例如“法国”到“巴黎”)是否会在最终标记位置以相同的深度产生因果作用。
Using four complementary causal diagnostics across four decoder-only models and eight prompt families, we find a robust temporal asymmetry: relation information becomes generation-controlling before entity information does.
通过在四种仅解码器(decoder-only)模型和八个提示词系列上使用四种互补的因果诊断方法,我们发现了一种稳健的时间不对称性:关系信息在实体信息之前就已开始控制生成过程。
Relation onset precedes entity onset by 10-16 tested layers (31-44% of network depth) at threshold 0.4, with the ordering holding across all 16 model-threshold combinations for thresholds 0.2-0.5.
在阈值为 0.4 时,关系信息的出现比实体信息早 10-16 个测试层(占网络深度的 31-44%),且在阈值 0.2-0.5 的所有 16 种模型-阈值组合中,这种先后顺序均保持一致。
Critically, entity information is not absent early: entity-token patching succeeds at 90-100% in early layers. Instead, entity commitment to generation is deferred: entity information is available at the entity-token position but becomes generation-controlling at the final token only after being routed there.
关键在于,实体信息并非在早期缺失:在早期层中,实体标记修补(entity-token patching)的成功率高达 90-100%。相反,实体对生成的承诺是被延迟的:实体信息在实体标记位置处是可用的,但只有在被路由到最终标记位置后,才会开始控制生成。