Simple Language Normalization Wins: Cross-Lingual Speaker Verification for the TidyVoice 2026 Challenge
Simple Language Normalization Wins: Cross-Lingual Speaker Verification for the TidyVoice 2026 Challenge
简单语言归一化胜出:TidyVoice 2026 挑战赛中的跨语言说话人验证
Abstract: Cross-lingual mismatch remains a key source of overall degradation in modern speaker verification. The TidyVoice2026 Challenge targets this setting with text-independent verification, comprising 3,666 training and 808 development speakers in 40 languages and 2,200 evaluation speakers in 38 unseen languages, without language labels at test time.
摘要: 跨语言不匹配仍然是现代说话人验证性能下降的主要原因之一。TidyVoice 2026 挑战赛针对这一场景进行了与文本无关的验证测试,数据集包含 40 种语言下的 3,666 名训练说话人和 808 名开发说话人,以及 38 种未见语言下的 2,200 名评估说话人,且在测试阶段不提供语言标签。
Starting from the official SimAM-ResNet34 baseline pretrained on VoxBlink2 and VoxCeleb2 and fine-tuned on TidyVoice, we revisit Nuisance Attribute Projection (NAP) as a simple language-normalization step in the embedding space.
我们以在 VoxBlink2 和 VoxCeleb2 上预训练并在 TidyVoice 上微调的官方 SimAM-ResNet34 基线模型为起点,重新审视了干扰属性投影(Nuisance Attribute Projection, NAP),将其作为嵌入空间中一种简单的语言归一化步骤。
We estimate a compact language subspace from cross-language same-speaker differences and project embeddings onto its orthogonal complement before cosine scoring with Adaptive Symmetric score normalization.
我们通过跨语言的同一说话人差异来估计一个紧凑的语言子空间,并在使用自适应对称(Adaptive Symmetric)分数归一化进行余弦评分之前,将嵌入向量投影到该子空间的补空间上。
This reduces development EER from 2.97% with cosine and 2.70% with AS-Norm to 2.18% and yields a Codabench evaluation score of 8.40, showing that simple back-end language normalization can rival more complex systems.
这一方法将开发集的等错误率(EER)从余弦评分的 2.97% 和 AS-Norm 的 2.70% 降低到了 2.18%,并在 Codabench 评估中获得了 8.40 的分数,这表明简单的后端语言归一化方法完全可以与更复杂的系统相媲美。