Metonymic Circuits for Abstract Concept Grounding in Vision Transformers
Computer Science > Artificial Intelligence arXiv:2610.06928 (cs) [Submitted on 3 Oct 2026] Title: Metonymic Circuits for Abstract Concept Grounding in Vision Transformers Authors: Jing Ding, Ziqiao Ma, Jiayuan Mao, Joyce Chai, Freda Shi.
计算机科学 > 人工智能 arXiv:2610.06928 (cs) [提交于 2026 年 10 月 3 日] 标题:视觉 Transformer 中用于抽象概念落地的转喻回路 作者:Jing Ding, Ziqiao Ma, Jiayuan Mao, Joyce Chai, Freda Shi。
Abstract: We study how Vision Transformers ground abstract concepts (e.g., angry) when training data provide limited direct referential evidence. We hypothesize a metonymic grounding mechanism in which abstract predictions are driven by concrete, interpretable anchor concepts (e.g., fire) that bridge visual signals to abstract semantics.
摘要:我们研究了当训练数据提供的直接指涉证据有限时,视觉 Transformer 如何将抽象概念(例如“愤怒”)落地。我们假设了一种转喻落地机制,其中抽象预测由具体的、可解释的锚点概念(例如“火”)驱动,这些概念将视觉信号与抽象语义联系起来。
By applying Transcoders on CLIP and DINO vision encoders, we recover intermediate features that can be associated with semantic labels for more concrete concepts, and trace their contributions in circuits underlying abstract concept recognition.
通过在 CLIP 和 DINO 视觉编码器上应用转码器(Transcoders),我们恢复了可以与更具体概念的语义标签相关联的中间特征,并追踪了它们在抽象概念识别底层回路中的贡献。
Experiments on a carefully curated icon dataset reveal structured metonymic circuits, in which perceptual primitives dominate early layers and object-like anchors precede abstract targets. Images containing rendered text instead recruit a distinct perceptual-to-textual route.
在精心策划的图标数据集上进行的实验揭示了结构化的转喻回路,其中感知基元在早期层中占主导地位,而类对象锚点先于抽象目标出现。包含渲染文本的图像则会调用一条独特的“感知到文本”路径。
Causal interventions further validate that metonymic intermediates are functionally involved in grounding abstract concepts.
因果干预进一步验证了转喻中间体在功能上参与了抽象概念的落地。