Relation Before Entity: Deferred Commitment in Language Model Factual Recall

Relation Before Entity: Deferred Commitment in Language Model Factual Recall

关系先于实体:语言模型事实回忆中的延迟承诺

Abstract: We ask whether relation-type information (e.g., capital-of) and entity-specific information (e.g., France to Paris) become causally active at the final-token position at the same depth during recall.

摘要: 我们探讨了在事实回忆过程中,关系类型信息(例如“……的首都”)和实体特定信息(例如“法国”到“巴黎”)是否会在最终标记位置以相同的深度产生因果作用。

Using four complementary causal diagnostics across four decoder-only models and eight prompt families, we find a robust temporal asymmetry: relation information becomes generation-controlling before entity information does.

通过在四种仅解码器(decoder-only)模型和八个提示词系列上使用四种互补的因果诊断方法,我们发现了一种稳健的时间不对称性:关系信息在实体信息之前就已开始控制生成过程。

Relation onset precedes entity onset by 10-16 tested layers (31-44% of network depth) at threshold 0.4, with the ordering holding across all 16 model-threshold combinations for thresholds 0.2-0.5.

在阈值为 0.4 时,关系信息的出现比实体信息早 10-16 个测试层(占网络深度的 31-44%),且在阈值 0.2-0.5 的所有 16 种模型-阈值组合中,这种先后顺序均保持一致。

Critically, entity information is not absent early: entity-token patching succeeds at 90-100% in early layers. Instead, entity commitment to generation is deferred: entity information is available at the entity-token position but becomes generation-controlling at the final token only after being routed there.

关键在于,实体信息并非在早期缺失:在早期层中,实体标记修补(entity-token patching)的成功率高达 90-100%。相反,实体对生成的承诺是被延迟的:实体信息在实体标记位置处是可用的,但只有在被路由到最终标记位置后,才会开始控制生成。