Verbalizable Representations Form a Global Workspace in Language Models

Verbalizable Representations Form a Global Workspace in Language Models

可言说表征在大语言模型中构成了全局工作空间

Abstract: Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible reasoning. In this paper, we present evidence that an analogous functional distinction has emerged in large language models. 摘要: 在人脑处理的所有信息中,只有极小一部分是“意识可及”的,即能够被口头报告、进行审慎控制以及灵活推理。在本文中,我们提供的证据表明,大语言模型中也出现了类似的函数功能区分。

Using a new interpretability technique, the Jacobian lens, we identify the representations a model is poised to verbalize at any point in its processing. These representations, which we collectively call the J-space, exhibit the functional properties characteristic of a global workspace: their contents can be reported, deliberately summoned and held, used to carry the intermediate steps of silent reasoning, and passed as arguments to arbitrary downstream computations, while automatic processing such as text parsing and routine inference proceeds without them. 通过一种新的可解释性技术——“雅可比透镜”(Jacobian lens),我们识别出了模型在处理过程中随时准备“言说”的表征。我们将这些表征统称为“J-空间”(J-space),它们展现出了全局工作空间(Global Workspace)的典型功能特性:其内容可以被报告、被刻意调取并保持、用于承载静默推理的中间步骤,并作为参数传递给任意下游计算;与此同时,文本解析和常规推理等自动处理过程则无需这些表征参与。

The J-space also has structural signatures that global workspace theory associates with conscious access: it carries coherent content only in an intermediate band of layers, holds on the order of tens of concepts at a time, and is broadcast by the model’s weights more widely than other representations. These properties make it a practical window into a model’s unspoken thinking. J-空间还具有全局工作空间理论中与“意识访问”相关的结构特征:它仅在中间层级中携带连贯的内容,一次能容纳数十个概念,并且通过模型权重进行的广播范围比其他表征更广。这些特性使其成为洞察模型“未言之思”的实用窗口。

In alignment audits, it reveals strategic deliberation, evaluation awareness, and trained-in misaligned dispositions that never appear in the model’s outputs. We find that post-training installs the Assistant’s point of view in the workspace, and we introduce counterfactual reflection training, which improves behavior by training only what a model would say if interrupted and asked to reflect. 在对齐审计中,它揭示了模型输出中从未显现的战略性审议、评估意识以及训练中植入的对齐偏差倾向。我们发现,后训练(post-training)过程将“助手视角”植入到了工作空间中;同时,我们引入了“反事实反思训练”(counterfactual reflection training),通过仅训练模型在被中断并要求反思时会说出的内容,从而改善其行为表现。

These results indicate that language models maintain a small, privileged set of representations bearing some of the functional hallmarks of conscious access, and that decoding these representations sheds light on ongoing cognitive processes. 这些结果表明,语言模型维护着一小部分特殊的表征集合,这些表征具备了意识访问的部分功能特征;解码这些表征,有助于我们深入理解模型正在进行的认知过程。