FD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech

FD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech

FD-VAD:用于流式全双工语音的语义端点检测

Abstract: Natural turn-taking in full-duplex voice interaction requires determining from partial speech whether a pause reflects hesitation or a completed conversational intent. Acoustic voice activity detection lacks this semantic information, while cascaded ASR-based endpointing introduces transcription dependence and additional processing stages.

摘要: 在全双工语音交互中,自然的轮流对话需要从部分语音中判断停顿是代表犹豫还是表达了完整的对话意图。传统的声学语音活动检测(VAD)缺乏这种语义信息,而基于级联自动语音识别(ASR)的端点检测则引入了对转录的依赖以及额外的处理阶段。

We formulate semantic endpoint detection as a causal audio-language reasoning task and introduce FD-VAD, an ASR-free streaming endpointer that maps bounded causal audio windows directly to Continue/Stop decisions. FD-VAD combines a frozen speech encoder with a lightweight modality adapter and a parameter-efficiently adapted language model, using a last-chunk training objective for streaming inference.

我们将语义端点检测建模为一种因果音频-语言推理任务,并提出了 FD-VAD。这是一种无需 ASR 的流式端点检测器,能够将有界的因果音频窗口直接映射为“继续”或“停止”的决策。FD-VAD 将冻结的语音编码器与轻量级模态适配器及参数高效适配的语言模型相结合,并采用最后分块(last-chunk)训练目标来实现流式推理。

We further introduce confidence-gated endpoint commitment to control interruption versus delay and boundary-focused hard-negative sampling to improve decisions around ambiguous turn boundaries. Across in-domain and conversational evaluations, FD-VAD outperforms strong streaming and non-streaming semantic turn classifiers, and achieves the highest EOT recall among qualifying systems on TurnBench dev set 0.853 (at FP<=0.10) in a zero-shot setting.

我们进一步引入了置信度门控端点提交机制来平衡中断与延迟,并采用聚焦于边界的难负样本采样来改善在模糊轮次边界处的决策表现。在领域内和对话评估中,FD-VAD 的表现优于强大的流式和非流式语义轮次分类器,并在 TurnBench 开发集上以零样本设置达到了 0.853 的 EOT(对话结束)召回率(在 FP<=0.10 时),在所有合格系统中表现最优。

These results show that semantic endpointing can be performed directly from streaming audio without intermediate ASR or dialogue state tracking.

这些结果表明,语义端点检测可以直接从流式音频中执行,而无需中间的 ASR 或对话状态跟踪。