A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID

A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID

冻结 BERT 中 AI 文本检测神经元的机制研究:基于 RAID 的稀疏探测与激活修补

Abstract: AI-generated text detectors achieve high accuracy on standard benchmarks, yet the internal representations that drive these predictions remain poorly understood. We study which neurons in a frozen BERT-base-uncased encoder support AI-text detection, using the RAID benchmark across six generators spanning pure-base and instruction-tuned models.

摘要: AI 生成文本检测器在标准基准测试中取得了很高的准确率,但驱动这些预测的内部表征仍未得到充分理解。我们研究了冻结的 BERT-base-uncased 编码器中哪些神经元支持 AI 文本检测,并使用了涵盖纯基础模型和指令微调模型的六种生成器的 RAID 基准进行测试。

We apply the L1-to-L2 sparse-probing protocol of Gurnee et al. (2023) to all 9,216 CLS hidden-state dimensions (12 layers x 768), which we call neurons. The procedure recovers a stable set of under 1% of neurons per generator, consistent across folds and seeds; a probe restricted to that set retains most of the full-feature detection accuracy.

我们将 Gurnee 等人 (2023) 的 L1-to-L2 稀疏探测协议应用于所有 9,216 个 CLS 隐藏状态维度(12 层 x 768,我们称之为神经元)。该过程为每个生成器恢复了一组不到 1% 的稳定神经元集合,且在不同折叠(folds)和随机种子下保持一致;仅限于该集合的探测器保留了大部分全特征检测的准确性。

Bidirectional activation patching confirms this set’s causal relevance: in both directions it flips predictions an order of magnitude more often than size-matched random sets. Mean-ablating the same neurons leaves accuracy largely intact; the signal is therefore redundantly distributed.

双向激活修补(Activation patching)证实了该集合的因果相关性:在两个方向上,它改变预测结果的频率比大小匹配的随机集合高出一个数量级。对相同神经元进行均值消融(Mean-ablating)后,准确率基本保持不变;因此,该信号是以冗余方式分布的。

Cross-generator analysis reveals a bipartite structure: instruction-tuned generators concentrate 30-36% of stable neurons in BERT’s final layer while both base generators fall below 14%, consistent with a layer-12 footprint of post-training alignment.

跨生成器分析揭示了一种二分结构:指令微调生成器将 30-36% 的稳定神经元集中在 BERT 的最后一层,而两个基础生成器均低于 14%,这与训练后对齐(post-training alignment)在第 12 层的特征印记相一致。

Leave-one-family-out evaluation shows the selected neurons retain 86-94% of the full-feature ceiling on unseen generator families, so a detector can operate on a small fixed subspace without re-identifying neurons per generator.

留一族评估(Leave-one-family-out evaluation)表明,所选神经元在未见过的生成器家族上保留了 86-94% 的全特征性能上限,因此检测器可以在一个小的固定子空间上运行,而无需为每个生成器重新识别神经元。