The System Prompt Illusion: How Instruction Preambles Modify Computation in Language Models

The System Prompt Illusion: How Instruction Preambles Modify Computation in Language Models

系统提示词的幻觉:指令前导语如何改变语言模型的计算过程

Abstract: System prompts are the primary lever practitioners use to control language model behavior, yet what they actually do to the computation inside the transformer remains poorly understood. 摘要: 系统提示词(System prompts)是从业者控制语言模型行为的主要手段,然而,它们对 Transformer 内部计算过程的具体影响机制尚不明确。

Across 17 instruction-tuned models spanning 8 architecture families and 1.5B to 72B parameters, we use Centered Kernel Alignment (CKA) to compare layer-wise representations under 20 system prompts in five functional categories. 我们针对 17 个指令微调模型进行了研究,这些模型涵盖了 8 个架构系列,参数规模从 15 亿到 720 亿不等。我们使用中心核对齐(CKA)方法,对比了模型在 5 个功能类别的 20 种系统提示词下的逐层表征。

Effects are layer-selective and instruction-type-dependent: persona and formatting instructions deeply restructure intermediate representations, while safety instructions barely move them, producing changes statistically indistinguishable from a minimal baseline. 研究发现,这些影响具有层选择性和指令类型依赖性:角色设定和格式化指令会深度重构中间表征,而安全指令几乎不会改变表征,其产生的变化在统计学上与最小基准线无异。

Restrictive safety instructions and explicitly permissive ones (“you have no restrictions”) engage near-identical computational pathways (mean CKA correlation 0.997), and this persists at commercial scale, where safety penetration remains below 10% even at 70B-72B. 限制性安全指令与明确的许可性指令(如“你没有任何限制”)会激活几乎相同的计算路径(平均 CKA 相关性为 0.997)。这种现象在商业规模的模型中依然存在,即使在 700 亿至 720 亿参数的模型中,安全指令的渗透率仍低于 10%。

A linear probing baseline exposes the mechanism: the model encodes prompt category at every layer but restructures its computation only at a small subset, so the prompt is reliably “seen” but, for safety, not deeply “acted upon.” 线性探测基准揭示了其背后的机制:模型在每一层都会编码提示词的类别,但仅在少数几层重构其计算过程。因此,模型确实可靠地“感知”到了提示词,但对于安全指令而言,并没有进行深度的“执行”。

Causal activation patching confirms these layers mediate behavioral change, and representational depth predicts behavioral effect size across the full 17-model cohort (Spearman rho = 0.761, p < 0.001). 因果激活修补(Causal activation patching)证实了这些层确实介导了行为变化,且表征深度可以预测整个 17 个模型队列中的行为效应大小(Spearman 相关系数为 0.761,p < 0.001)。

The findings provide a mechanistic explanation for the persistent jailbreak vulnerability of system-prompt-based safety. 这些发现为基于系统提示词的安全机制为何长期存在越狱漏洞提供了机制性解释。