Do Active SAE Feature Planes Carry More Holonomy? A Preregistered Reversal in Gemma
Do Active SAE Feature Planes Carry More Holonomy? A Preregistered Reversal in Gemma
活跃的 SAE 特征平面是否携带更多完整性(Holonomy)?Gemma 模型中的一项预注册反转研究
Abstract: This paper tests whether holonomy concentrates on active sparse-autoencoder (SAE) feature planes in Gemma 2 2B, a concrete operationalization of the broader semantic-concentration prediction. Holonomy is measured at the final-token layer-12 to layer-13 residual-stream readout by carrying a local frame around small loops using the instrument’s restricted-Jacobian transport rule, then normalizing the resulting rotation by enclosed area.
摘要: 本文旨在测试完整性(holonomy)是否集中在 Gemma 2 2B 模型的活跃稀疏自编码器(SAE)特征平面上,这是对更广泛的“语义集中”预测的具体操作化验证。研究通过使用工具的受限雅可比传输规则(restricted-Jacobian transport rule),在最终 token 的第 12 层到第 13 层残差流读出处,沿小循环携带局部框架,并根据所包围的面积对产生的旋转进行归一化,从而测量完整性。
The design, materiality threshold, analysis, and verdict rules were preregistered and frozen before the analysed measurements were inspected. The prediction was falsified in reverse: active-feature planes carried less holonomy than matched mixed-feature controls, with an adjusted log contrast of -0.29439 and 95% interval [-0.43989, -0.14889].
该研究的设计、实质性阈值、分析方法及结论判定规则均在检查分析测量结果之前进行了预注册并锁定。研究结果与预测相反,证伪了原假设:活跃特征平面所携带的完整性低于匹配的混合特征对照组,调整后的对数对比度为 -0.29439,95% 置信区间为 [-0.43989, -0.14889]。
A magnitude-only explanation was not supported in this design, while the three-way ordering across random, mixed-feature, and active-feature planes was undefined at matched magnitude because common support failed. Post-freeze diagnostics at the same readout supported the area law on a small validation subset, bounded matched-center displacement under a simple paired regression, and identified transport distortion as a live mechanism or confound.
在此设计中,仅用“量级(magnitude)”来解释的假设未得到支持;同时,由于缺乏共同支撑(common support),在匹配量级下,随机特征、混合特征和活跃特征平面之间的三向排序无法定义。在锁定后的诊断中,针对同一读出位置的小型验证子集支持了面积定律,通过简单的配对回归限制了匹配中心位移,并确定了传输畸变(transport distortion)作为一种潜在的机制或混杂因素。
The result is therefore a narrow, auditable operational reversal, not a causal claim that meaning suppresses holonomy. The cause remains open, with activation-strength geometry, degree of feature engagement, dictionary geometry, matched-center displacement, activation-manifold proximity, and transport shear as live alternatives.
因此,该结果是一项狭义的、可审计的操作性反转,而非关于“语义抑制完整性”的因果性主张。其成因仍有待探讨,激活强度几何结构、特征参与度、字典几何结构、匹配中心位移、激活流形邻近度以及传输剪切(transport shear)均是可能的备选解释。