Capacity, Responsiveness and Alignment: What Makes a Latent Structure Actionable

Localizing latent structures in the activation space of language models (LMs) is central to understanding and controlling their behavior. Yet, localized structures can differ substantially in their causal influence, raising the question of what makes a structure actionable.

在语言模型(LM)的激活空间中定位潜在结构,对于理解和控制其行为至关重要。然而,已定位的结构在因果影响方面可能存在显著差异,这就提出了一个问题:究竟是什么使得一个结构具有可操作性。

We tackle this question by casting causal influence as a product of three factors and showing empirically that they act as interpretable, distinct constraints: capacity, measuring the sensitivity of the model’s output to movement along the structure, responsiveness, capturing how promotable the concept is given the current context, and alignment, reflecting how well the structure aligns with the context-specific representation of the concept.

我们通过将因果影响视为三个因素的乘积来解决这个问题,并从实证角度证明它们作为可解释且独特的约束发挥作用:容量(capacity),衡量模型输出对沿结构移动的敏感度;响应性(responsiveness),捕捉在当前语境下该概念的可提升程度;以及对齐(alignment),反映结构与该概念在特定语境下的表示的契合程度。

Across 4 LM families and 50 concepts, we observe that causal effectiveness requires all factors to be high; low capacity and responsiveness reduce it by 84% and 95%, respectively, while low alignment can reverse it, suppressing concept expression.

通过对 4 个语言模型系列和 50 个概念的研究,我们观察到因果有效性要求所有因素都保持高水平;低容量和低响应性分别使其降低了 84% 和 95%,而低对齐度则可能产生反向效果,从而抑制概念的表达。

Moreover, we find that causality is context-dependent rather than an intrinsic property of the structure, with causally effective directions forming a low-dimensional subspace that varies across contexts. By restricting the training of linear probes to this subspace, we introduce causal probes that achieve 17%-118% improvement in steering across models, with only 3% reduction in concept detection.

此外,我们发现因果关系是依赖于语境的,而非结构的固有属性,具有因果有效性的方向形成了一个随语境变化的低维子空间。通过将线性探针的训练限制在该子空间内,我们引入了因果探针,在跨模型引导方面实现了 17% 到 118% 的提升,而概念检测能力仅下降了 3%。