Tokka-Bench: Evaluating Tokenizers Across 100 Natural and 20 Programming Languages
Computer Science > Computation and Language arXiv:2610.08794 (cs) [Submitted on 25 Mar 2026] Title: Tokka-Bench: Evaluating Tokenizers Across 100 Natural and 20 Programming Languages Authors: Ben Gubler.
计算机科学 > 计算与语言 arXiv:2610.08794 (cs) [提交于 2026 年 3 月 25 日] 标题:Tokka-Bench:评估跨越 100 种自然语言和 20 种编程语言的分词器 作者:Ben Gubler。
Abstract: Large language models rely on subword tokenizers whose quality varies across languages, yet no standardized multi-metric framework exists for broad comparative evaluation. We introduce Tokka-Bench, an open-source framework that evaluates tokenizers on five complementary metrics — bytes per token, unique token coverage, subword fertility, word-split rate, and vocabulary composition — across 100 natural languages (30+ scripts) and 20 programming languages, using language-aware segmentation adapted to each writing system.
摘要:大型语言模型依赖于子词分词器,其质量在不同语言间存在差异,但目前尚无标准化的多指标框架用于广泛的比较评估。我们推出了 Tokka-Bench,这是一个开源框架,通过五种互补指标——每个标记的字节数、唯一标记覆盖率、子词丰富度、单词拆分率和词汇组成——来评估分词器,涵盖 100 种自然语言(30 多种书写系统)和 20 种编程语言,并针对每种书写系统采用了语言感知分段技术。
Comparing seven BPE tokenizers (GPT-2, GPT-4, gpt-oss, Llama 3.1, Gemma 3, Qwen3, and Kimi K2) within individual languages, we find that vocabulary allocation strategy matters more than raw vocabulary size, and that programming-language efficiency has converged among recent tokenizers despite divergent natural-language profiles. The framework, data, and interactive dashboard are publicly available.
通过比较七种 BPE 分词器(GPT-2、GPT-4、gpt-oss、Llama 3.1、Gemma 3、Qwen3 和 Kimi K2)在各语言中的表现,我们发现词汇分配策略比原始词汇量更为重要,且尽管自然语言特征各异,近期分词器在编程语言处理效率上已趋于一致。该框架、数据和交互式仪表板均已公开。