Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability

Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability

机械层析成像:面向控制可解释性的设计测量

Abstract: Mechanistic interpretability seeks quantities that models do not expose directly: represented states, component effects, interactions, and responses to interventions. Patching, gradients, Hessian-vector products, and subset interventions provide different measurements under different access assumptions and may target different quantities.

摘要: 机械可解释性旨在寻求模型无法直接暴露的量:表征状态、组件效应、交互作用以及对干预的响应。修补(Patching)、梯度、Hessian 向量积和子集干预在不同的访问假设下提供不同的测量方法,并可能针对不同的目标量。

We formulate their shared measurement structure as mechanistic tomography: designed measurement for recovering internal mechanisms and intervention effects. For a chosen basis and intervention family, measurements take the form y = Ax + w, where A describes the interventions, x is the target map, and w contains nonlinear response, sampling error, and basis misspecification.

我们将它们共有的测量结构表述为“机械层析成像”:即为恢复内部机制和干预效应而设计的测量。对于选定的基底和干预族,测量形式为 y = Ax + w,其中 A 描述干预,x 是目标映射,w 包含非线性响应、采样误差和基底设定偏差。

This language gives a practical procedure: start with the least costly measurements, test on held-out interventions at the intended scale, calibrate simple mismatch, and expand the measurement family when structured residuals remain. Control provides a demanding validation setting because an estimate that guides an intervention acts as an observer. In a two-HMM model, control error rises with observer error, while target improvement can hide nuisance-state movement.

这种语言提供了一种实用的流程:从成本最低的测量开始,在预期的规模上对留出的干预进行测试,校准简单的失配,并在存在结构化残差时扩展测量族。控制提供了一个严苛的验证环境,因为引导干预的估计值充当了观察者的角色。在双隐马尔可夫模型(two-HMM model)中,控制误差会随观察者误差而上升,而目标改进可能会掩盖干扰状态的变动。

Under forward-only access, sparse aggregate measurements recover a finite-effect map with fewer interventions than coordinate patching. With gradient access, finite probes improve a local attribution map. Lifted measurements and Hessian-vector products recover interactions missed by first-order maps, while Tracr shows that the required family depends on the basis.

在仅前向访问的情况下,稀疏聚合测量比坐标修补(coordinate patching)能以更少的干预次数恢复有限效应映射。在具有梯度访问权限时,有限探测可以改进局部归因映射。提升测量(Lifted measurements)和 Hessian 向量积可以恢复一阶映射所遗漏的交互作用,而 Tracr 表明所需的测量族取决于基底的选择。

On GPT-2-small IOI, the Name Mover-Negative Name Mover interaction is the largest held-out predictive term among three tested cross-group pairs. On Qwen-2.5-7B, finite calibration makes an additive refusal-response map adequate, so held-out error does not support pairwise lifting.

在 GPT-2-small 的 IOI 任务上,“Name Mover-Negative Name Mover”交互作用是三个测试的跨组对中最大的留出预测项。在 Qwen-2.5-7B 上,有限校准使得加性拒绝响应映射(additive refusal-response map)已足够充分,因此留出误差不支持成对提升(pairwise lifting)。