Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models

Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models

语音与面部表情中的情感:多模态基础模型中共享的情感机制

Abstract: Modern multimodal foundation models (MFMs) have made rapid progress on tasks requiring integrated perception across speech, vision, and language, including emotion recognition. However, it remains unclear whether they recognize speech and facial emotion through shared affective functional units or modality-specific pathways.

摘要: 现代多模态基础模型(MFMs)在需要跨语音、视觉和语言进行综合感知的任务(包括情感识别)中取得了飞速进展。然而,目前尚不清楚它们是通过共享的情感功能单元,还是通过特定模态的路径来识别语音和面部情感的。

We explore emotion-sensitive neurons (ESNs), sparse decoder neurons selectively associated with emotion categories, in three MFMs: Gemma-4-12B-it, MiniCPM-o-4.5, and Qwen2.5-Omni-7B. Using speech emotion recognition and facial expression recognition as complementary probes, we identify acoustic and visual ESNs.

我们探索了三种多模态基础模型(Gemma-4-12B-it、MiniCPM-o-4.5 和 Qwen2.5-Omni-7B)中的“情感敏感神经元”(ESNs),即与特定情感类别选择性关联的稀疏解码器神经元。通过将语音情感识别和面部表情识别作为互补的探测手段,我们识别出了声学和视觉 ESN。

Visual ESNs are causally meaningful: deactivating them selectively impairs recognition of the associated facial emotion, whereas steering their activations selectively enhances recognition of that emotion relative to other emotion categories. Acoustic and visual ESNs further show emotion-matched overlap and similar layer-wise distributions, indicating partial structural alignment between affective representations across speech and faces.

视觉 ESN 具有因果意义:停用它们会选择性地削弱对相关面部情感的识别,而引导它们的激活则会相对于其他情感类别,选择性地增强对该情感的识别。声学和视觉 ESN 进一步显示出情感匹配的重叠和相似的层级分布,这表明语音和面部情感表征之间存在部分结构对齐。

Finally, cross-modal interventions reveal bidirectional causal transfer: ESNs identified from one modality produce emotion-specific effects when applied to the other. Our findings provide one of the first cross-modality activation-level analyses of affective functional units in MFMs, suggesting that speech and facial emotion recognition partially converge onto sparse decoder-level components that can be localized and manipulated without training.

最后,跨模态干预揭示了双向的因果迁移:从一种模态中识别出的 ESN 在应用于另一种模态时,会产生特定于情感的效果。我们的研究结果提供了多模态基础模型中情感功能单元的首批跨模态激活水平分析之一,表明语音和面部情感识别部分汇聚于稀疏解码器层级的组件上,且这些组件无需训练即可被定位和操纵。