Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models

Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models

区分语音语言模型中的决策规则失调与读出覆盖局限性

Abstract: Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures at different stages of the audio-to-answer computation. 摘要: 语音语言模型在副语言任务中的评估日益依赖于提示回答的准确性,但回答的准确性往往混杂了从音频到回答计算过程中不同阶段的失败。

We introduce a generation-aligned diagnostic ladder that compares the emitted answer, the option logits, an affine readout of those logits, and a linear readout of the hidden state at the same answer token. Successive differences separate endpoint, decision-rule, and readout-coverage gaps. 我们引入了一种生成对齐的诊断阶梯,通过比较生成的回答、选项逻辑值(logits)、这些逻辑值的仿射读出,以及同一回答标记处隐藏状态的线性读出,来分析模型表现。通过逐级差异分析,我们成功区分了端点差距、决策规则差距和读出覆盖差距。

Across five systems and two emotion corpora, state decoding exceeds generation by 27.8 accuracy points on average, and both the decision-rule and readout-coverage gaps are positive in all ten conditions. A label-free logit correction improves generated accuracy in every condition, showing that part of the decision-rule gap is actionable. 在五个系统和两个情感语料库的测试中,状态解码的平均准确率比生成高出 27.8 个百分点,且在所有十种实验条件下,决策规则差距和读出覆盖差距均为正值。一种无需标签的逻辑值校正方法在所有条件下都提升了生成准确率,这表明决策规则差距中的一部分是可以被优化的。

In rank-matched comparisons, emotion information outside the native readout generalizes to held-out speakers and survives controls for measured acoustic descriptors, but replacing the selected readout-external directions usually has little effect on emitted answers. These results distinguish information availability from behavioral use and localize performance losses across the decision rule and the state-to-answer readout. 在秩匹配比较中,原生读出之外的情感信息能够泛化至留存的说话人,且在控制了测量声学描述符后依然存在;然而,替换所选的读出外部方向通常对生成的回答影响甚微。这些结果区分了信息的可用性与行为上的使用,并将性能损失定位在决策规则和状态到回答的读出过程之中。