The Price of Token Boundaries: Compression Certificates and Prediction

The Price of Token Boundaries: Compression Certificates and Prediction

代币边界的代价:压缩证明与预测

Abstract: Pre-tokenisation restricts which text fragments can become prediction units, but its compression cost is obscured when tokenisers are compared only under the same boundaries. We measure this cost by bounding the minimum token count from both sides, with and without a regular-expression boundary rule. Nonnegative prices on token occurrences yield a lower bound through shortest paths and vocabulary-budget selection; maximising over all prices recovers the linear programming relaxation, and an independent integer checker certifies the reported values.

摘要: 预分词(Pre-tokenisation)限制了哪些文本片段可以成为预测单元,但当分词器仅在相同边界下进行比较时,其压缩代价往往被掩盖。我们通过在有或无正则表达式边界规则的情况下,从两侧界定最小代币数量来衡量这一代价。通过最短路径和词汇预算选择,代币出现的非负价格产生了一个下界;对所有价格进行最大化处理可恢复线性规划松弛,且一个独立的整数检查器会对报告的数值进行验证。

On English Wikipedia, boundaries increase the optimal token count by 28.3—36.8%. Byte pair encoding lies 2.1% above the constrained lower bound, but 10.9% above the unrestricted bound. Compression and prediction favour different dictionaries: at 85M non-embedding parameters and matched training-token budgets, unrestricted fitting yields higher mean held-out bits per byte under a common unrestricted decoder in all 12 languages in the paired study and 11 of 12 under independent tuning and evaluation.

在英文维基百科上,边界使最优代币数量增加了 28.3%—36.8%。字节对编码(BPE)比受限下界高出 2.1%,但比无限制下界高出 10.9%。压缩和预测偏好不同的词典:在 85M 非嵌入参数和匹配训练代币预算的情况下,在配对研究的所有 12 种语言中,以及在独立调优和评估的 12 种语言中的 11 种里,无限制拟合在通用无限制解码器下产生了更高的平均留出比特每字节(bits per byte)。

To study intermediate boundary policies, we introduce boundary licences, which limit the vocabulary entries permitted to cross cuts and admit the same form of certificate. On separate English and Chinese fitting corpora, licensing 10% of the vocabulary budget recovers 85.2% and 100.0% of the achieved token-count reduction from removing all cuts. These results quantify the compression cost of boundaries while separating it from the prediction quality of the resulting token units.

为了研究中间边界策略,我们引入了“边界许可”(boundary licences),它限制了允许跨越切分的词汇条目,并支持相同形式的证明。在独立的英、中拟合语料库上,许可 10% 的词汇预算可以恢复移除所有切分所带来的代币数量减少量的 85.2% 和 100.0%。这些结果量化了边界的压缩代价,同时将其与所得代币单元的预测质量分离开来。