Represented but Ignored: A Causal Account of Prosodic Underuse in Audio-Language Models

Represented but Ignored: A Causal Account of Prosodic Underuse in Audio-Language Models

已表征却被忽视:音频语言模型中韵律利用不足的因果分析

Abstract: Human speech is richly expressive, with prosody carrying linguistic and emotional information beyond the lexical content. A capable large audio-language model (audio-LLM) should therefore support expressive speech understanding, not only transcribing what was said but also interpreting how it was said.

摘要: 人类语言具有丰富的表现力,韵律承载着词汇内容之外的语言和情感信息。因此,一个强大的大型音频语言模型(audio-LLM)应支持富有表现力的语音理解,不仅要转录“说了什么”,还要解读“是如何说的”。

Yet behavioral evaluations alone cannot reveal why a model fails on prosodic input. An error may reflect loss of acoustic information, incorrect internal interpretation, or failure to use a representation that is already available inside the model.

然而,仅靠行为评估无法揭示模型为何在处理韵律输入时失败。错误可能反映了声学信息的丢失、内部解读的错误,或是未能利用模型内部已经存在的表征。

We introduce a stage-specific probe ladder for localizing these failure modes in audio-LLMs. Across four understanding-only audio-LLMs, prosodic information is usually preserved in the audio path and decodable in late LLM states. Yet it is only partially expressed in the model’s final response.

我们引入了一种分阶段的探测阶梯(probe ladder),用于定位音频语言模型中的这些故障模式。在四个仅具备理解能力的音频语言模型中,韵律信息通常被保留在音频路径中,并可在大型语言模型的后期状态中解码。然而,这些信息在模型的最终回答中仅得到了部分体现。

We test the causal status of this latent representation with targeted hidden-state interventions. Every intervention shifts the answer distribution in the predicted direction, and in most model—task cells a single edit at the relevant layer is sufficient to drive the model toward the suppressed prosodic decision, though this recovery is directional rather than a selective restoration of the correct class.

我们通过针对性的隐藏状态干预,测试了这种潜在表征的因果地位。每一次干预都会使回答分布向预测方向偏移,并且在大多数模型-任务单元中,在相关层进行一次编辑就足以驱动模型做出被抑制的韵律决策,尽管这种恢复是方向性的,而非对正确类别的选择性还原。

Feature-level analysis further suggests that this recoverable signal can be expressed through a small subspace. Some of the highest-attribution features in this analysis align with acoustic cues known to carry prosodic information.

特征层面的分析进一步表明,这种可恢复的信号可以通过一个小的子空间来表达。在此分析中,一些归因度最高的特征与已知承载韵律信息的声学线索相吻合。

Within the matched-content contrasts we test, these results locate the recurring bottleneck not in perceiving prosody but in using it. Models that hear and correctly represent a prosodic cue can still fail to express it in their answers.

在我们测试的匹配内容对比中,这些结果表明反复出现的瓶颈不在于感知韵律,而在于利用韵律。能够听到并正确表征韵律线索的模型,在回答时仍可能无法将其表达出来。