Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning

Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning

Open TTS 排行榜:多语言语音合成与声音克隆的可扩展评估

The pace of open-source text-to-speech (TTS) model releases has been incredible. On the Hugging Face Hub (as of Sep 30, 2026) there are more than 8K TTS models available 🚀 Evaluation, however, hasn’t kept pace: it remains fragmented and unstandardized. 开源语音合成(TTS)模型的发布速度令人惊叹。截至 2026 年 9 月 30 日,Hugging Face Hub 上已有超过 8,000 个 TTS 模型可用 🚀 然而,评估工作却未能跟上步伐:它仍然处于碎片化且缺乏标准化的状态。

The gold standard is human preference scores such as MOS or MUSHRA. To this end, several arena-based leaderboards have established themselves as useful reference points for the community: TTS Arena v2, Artificial Analysis, and Voice Arena. These arenas compare models by presenting users with TTS outputs from two models, and asking them to choose one over the other. After collecting a sufficient number of votes, an Elo score is computed to rank models, typically with the Bradley–Terry model. 黄金标准是诸如 MOS 或 MUSHRA 等人类偏好评分。为此,几个基于竞技场的排行榜已成为社区有用的参考点:TTS Arena v2、Artificial Analysis 和 Voice Arena。这些竞技场通过向用户展示两个模型的 TTS 输出并要求用户进行选择来比较模型。在收集到足够多的投票后,通常会使用 Bradley–Terry 模型计算 Elo 分数来对模型进行排名。

While human preference is the ultimate decider, arenas cannot scale to keep up with the pace of TTS releases. This may partly explain why open-source models are underrepresented on arena-style leaderboards: as of Sep 30, 2026, only 16 of the 92 models on Artificial Analysis are open-weights, with a similar skew on Voice Arena. This likely reflects practical factors: adding an API model requires little more than an API key, whereas an open model must be hosted and served by the arena operator, and commercial providers have more reason to seek placement than open-source authors. 虽然人类偏好是最终的决定因素,但竞技场无法扩展以跟上 TTS 发布的步伐。这在一定程度上解释了为什么开源模型在竞技场式排行榜中代表性不足:截至 2026 年 9 月 30 日,Artificial Analysis 上的 92 个模型中只有 16 个是开放权重的,Voice Arena 的情况也类似。这可能反映了实际因素:添加一个 API 模型只需一个 API 密钥,而开源模型必须由竞技场运营商进行托管和服务,且商业提供商比开源作者更有动力寻求排名。

Another limitation with arena-style evaluation is voter consistency: no arena can ensure that the same voters with the same criteria of “better” can consistently evaluate models over time. Even the preferences of a single person change over time (“A man cannot step into the same river twice” as famously said by Heraclitus). 竞技场式评估的另一个局限性是投票者的一致性:没有任何竞技场能确保同一批投票者在不同时间点能以相同的“更好”标准持续评估模型。即使是一个人的偏好也会随时间而改变(正如赫拉克利特所言:“人不能两次踏进同一条河流”)。

To this end, we’ve built the Open TTS Leaderboard, which uses objective metrics to evaluate models on complementary aspects of performance: 为此,我们构建了 Open TTS 排行榜,它使用客观指标从互补的性能维度评估模型:

  • Intelligibility: word/character error rate (WER and CER) between the prompt and the generated audio’s transcript, using Qwen3 ASR.
  • 可懂度: 使用 Qwen3 ASR 计算提示词与生成音频转录文本之间的词错误率/字符错误率(WER 和 CER)。
  • Speed: inverse real-time factor (RTFx) for batched offline inference on an H200 GPU, and time-to-first-audio (TTFA) for quantifying streaming batch size 1 latency on an H200 GPU and CPU.
  • 速度: 在 H200 GPU 上进行批量离线推理的逆实时因子(RTFx),以及用于量化 H200 GPU 和 CPU 上流式批处理大小为 1 的延迟的首音频时间(TTFA)。
  • Speaker similarity: by computing the cosine similarity (SIM) between WavLM speaker embeddings of the generated audio and the reference clip.
  • 说话人相似度: 通过计算生成音频与参考片段的 WavLM 说话人嵌入之间的余弦相似度(SIM)。

By relying on objective metrics, evaluating a model drops from a couple weeks (for collecting votes) to a couple hours ⚡ Importantly, the Open TTS Leaderboard does not replace human preference ranking. ASR-based WER provides a proxy for intelligibility, while speaker similarity estimates voice identity preservation. Neither directly measures naturalness, expressiveness, or listener preference. Nevertheless, they can even inform voting-based leaderboards which models to include in their evaluations. 通过依赖客观指标,评估一个模型所需的时间从几周(用于收集投票)缩短到了几个小时 ⚡ 重要的是,Open TTS 排行榜并不能取代人类偏好排名。基于 ASR 的 WER 提供了可懂度的代理指标,而说话人相似度则评估了声音身份的保留程度。两者都不能直接衡量自然度、表现力或听众偏好。尽管如此,它们可以为基于投票的排行榜提供参考,帮助其决定将哪些模型纳入评估。

Multilingual + voice cloning evaluation

多语言 + 声音克隆评估

From the default view of the leaderboard, models are ranked by macro-average WER on the English splits of Seed TTS Eval and CV3 Eval. hexgrad/Kokoro-82M, Supertone/supertonic-3, and fishaudio/s2-pro lead the pack on English WER when averaged on these two splits. English performance doesn’t necessarily translate to other languages. Multiple languages can be toggled to rank models on multilingual performance. 在排行榜的默认视图中,模型根据 Seed TTS Eval 和 CV3 Eval 的英语部分上的宏平均 WER 进行排名。在这些数据集的平均 WER 上,hexgrad/Kokoro-82M、Supertone/supertonic-3 和 fishaudio/s2-pro 处于领先地位。英语表现并不一定能转化为其他语言的表现。可以通过切换多种语言来对模型的多语言性能进行排名。

Note that Chinese, Japanese, and Korean are character-based languages and so character error rate (CER) is reported, and the “Average WER” across languages is a macro-average across languages. k2-fsa/OmniVoice, fishaudio/s2-pro, and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 are strong multilingual models. 请注意,中文、日语和韩语是基于字符的语言,因此报告的是字符错误率(CER),而跨语言的“平均 WER”是各语言的宏平均值。k2-fsa/OmniVoice、fishaudio/s2-pro 和 FunAudioLLM/Fun-CosyVoice3-0.5B-2512 是强大的多语言模型。

Compare and vote on TTS outputs

比较并投票 TTS 输出

Numbers only tell part of the story, and as mentioned earlier human preference is the ultimate decider. From the “Listen” tab, you can compare the generated outputs that are behind the metrics, to find which model(s) you prefer! The “Listen” tab fills an important gap in existing TTS leaderboards: a space to explore model outputs of various models. You can even give feedback on the generated outputs. 数字只能说明部分情况,正如前面提到的,人类偏好是最终的决定因素。在“Listen”(收听)选项卡中,您可以比较指标背后的生成输出,找到您更喜欢的模型!“Listen”选项卡填补了现有 TTS 排行榜的一个重要空白:一个探索各种模型输出的空间。您甚至可以对生成的输出提供反馈。

Streaming performance

流式性能

The “Streaming” tab compares the streaming capabilities. Models are ranked by TTFA (time-to-first-audio), which quantifies how long a user waits after probing a model in order to obtain audio that can be played. This is important for voice agents and other interactive applications. “Streaming”(流式)选项卡比较了流式传输能力。模型按 TTFA(首音频时间)排名,该指标量化了用户在请求模型后等待多久才能获得可播放音频。这对语音代理和其他交互式应用至关重要。