Representational Simplicity and Circuit Size Dissociate in a Threshold-Dependent Way: A Controlled Test via Adversarial Training
Representational Simplicity and Circuit Size Dissociate in a Threshold-Dependent Way: A Controlled Test via Adversarial Training
表征简洁性与电路规模在阈值依赖下的解耦:基于对抗训练的受控测试
Abstract: Sparse-autoencoder decomposability and concentrated feature attribution are increasingly treated as evidence that a model’s computation is easier to reverse-engineer. Whether this representational and attributional cleanliness actually predicts a smaller or more tractable causal circuit remains an open question. We test this directly using adversarial training as a controlled instrument: it reliably reshapes internal representations, but this alone does not constitute a test of circuit size.
摘要: 稀疏自编码器(Sparse-autoencoder)的可分解性和集中的特征归因正日益被视为模型计算更易于逆向工程的证据。这种表征和归因上的简洁性是否真的预示着更小或更易处理的因果电路,仍是一个悬而未决的问题。我们使用对抗训练作为受控工具对此进行了直接测试:它能可靠地重塑内部表征,但这本身并不构成对电路规模的测试。
We investigate this question through reverse-engineering complexity: the causal structure required to recover a model’s behavior at a fixed level of faithfulness. To our knowledge, this is the first controlled empirical test of whether representational or attributional simplicity translates into causal simplicity at the circuit level.
我们通过逆向工程复杂度来研究这一问题:即在固定的保真度水平下恢复模型行为所需的因果结构。据我们所知,这是首次关于表征或归因简洁性是否会转化为电路层面因果简洁性的受控实证测试。
Starting from the same pretrained GPT-2 Small checkpoint, we apply matched standard and adversarial continual training, requiring both conditions to retain competence on indirect object identification and pass independent robustness verification before comparing mechanisms. We then compare the models along three complementary axes: sparse-autoencoder decomposability, SAE feature engagement in task attribution, and the size of faithful circuits recovered from the raw computational graph.
我们从同一个预训练的 GPT-2 Small 检查点出发,应用匹配的标准训练和对抗性持续训练,要求两种条件在比较机制之前,都必须保持对间接宾语识别(IOI)任务的能力,并通过独立的鲁棒性验证。随后,我们从三个互补的维度对模型进行了比较:稀疏自编码器(SAE)的可分解性、SAE 特征在任务归因中的参与度,以及从原始计算图中恢复出的忠实电路的规模。
The robust model is more SAE-decomposable and engages fewer SAE features in task attribution. Circuit size is regime-dependent: on competence-matched IOI, standard leads or ties below 85% faithfulness, but robust needs substantially fewer edges at high faithfulness (90%, 95%), a pattern established on the primary pair while representational trends generalize across a seven-point sweep and a second corpus.
鲁棒模型具有更高的 SAE 可分解性,且在任务归因中参与的 SAE 特征更少。电路规模取决于具体区间:在能力匹配的 IOI 任务中,当保真度低于 85% 时,标准模型表现领先或持平;但在高保真度(90%、95%)下,鲁棒模型所需的边数显著更少。这一模式在主要实验对中确立,且表征趋势在七点扫描和第二个语料库中均表现出泛化性。