A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space: What the Units of Voynichese Are Not
A Glyph Is Not a Letter, a Token Is Not a Word, a Space Is Not a Space: What the Units of Voynichese Are Not
字形非字母,标记非单词,空格非空格:伏尼契文的单位究竟是什么?
Abstract: The Voynich manuscript (Beinecke MS 408) is usually analysed on three unstated assumptions: that its glyphs are letters, that the strings between blanks are words, and that every blank is a word space. We test all three against the Zandbergen-Landini transliteration with matched prose, cipher, and pseudo-text controls and quire-level resampling. None holds, and the failures share a shape: the order in Voynichese sits at the edges of tokens and at graded boundaries between them, not in the succession of tokens themselves.
摘要: 伏尼契手稿(Beinecke MS 408)的分析通常基于三个未明说的假设:其字形(glyphs)即字母,空格间的字符串即单词,且每个空格都是单词间距。我们利用 Zandbergen-Landini 转写方案,结合匹配的散文、密码文本、伪文本对照组以及手稿折叠单位(quire-level)重采样,对这三个假设进行了验证。结果显示,这三个假设均不成立,且其失效模式具有共同特征:伏尼契文的顺序逻辑存在于标记(tokens)的边缘及其间的渐变边界上,而非标记本身的先后排列中。
Glyph regularity is too strong for one-to-one substitution of any tested plaintext (conditional entropy 2.7 bits against about 3.5 for Latin, Italian, and English) and resolves instead onto a quire-stable scale of recurrent multi-symbol units. Tokens form a plausible vocabulary, yet the identity of one token predicts the next by under 1% of token entropy, below every matched control (2-10%), while the glyphs at token edges share 0.2 bits of mutual information, more than in any prose control.
字形的规律性过强,无法通过任何测试过的明文进行一一对应替换(条件熵为 2.7 比特,而拉丁语、意大利语和英语约为 3.5 比特),其规律反而体现为一种在手稿折叠单位间稳定的、重复出现的多符号单元尺度。虽然这些标记构成了一个看似合理的词汇表,但一个标记对下一个标记的预测能力不到标记熵的 1%,低于所有匹配的对照组(2-10%);与此同时,标记边缘的字形共享了 0.2 比特的互信息,这高于任何散文对照组。
Blanks fall into two regimes: the separators transcribers marked uncertain behave like word-internal junctures, are physically narrower on the page (AUC 0.905 from independent image coordinates, with the same sign in a small blind ink audit), and are crossed by learned units even when every space is erased before learning. This profile is also what discriminates. A published Voynich-imitating cipher and a self-citation text generator both reproduce the low entropy, the unit scale, the weak token order, and the null result of a calibrated substitution attack; neither reproduces the edge-glyph coupling or the open, hapax-rich vocabulary (70% singleton types against 41% and 59-60%).
空格分为两种情况:转录者标记为“不确定”的分隔符表现得更像是单词内部的连接点,在页面上物理宽度更窄(根据独立图像坐标计算的 AUC 为 0.905,在小规模盲测墨迹审计中结果一致),且即使在学习前抹除所有空格,这些连接点仍会被习得的单位跨越。这种特征分布也是区分的关键。已发表的伏尼契模仿密码和自引文本生成器都能重现低熵、单位尺度、弱标记顺序以及校准替换攻击的无效结果;但两者都无法重现边缘字形的耦合特征,也无法重现那种开放且包含大量单次出现词汇(hapax)的词汇表(单次出现类型占比 70%,而对照组为 41% 和 59-60%)。
Any account of the manuscript must therefore earn, rather than assume, the step from glyphs, tokens, and separators to letters, words, and word spaces, and these are the measurements on which to do so.
因此,任何关于该手稿的解释都必须通过论证而非假设,来完成从“字形、标记、分隔符”到“字母、单词、单词间距”的跨越,而本文所提供的测量数据正是实现这一论证的基础。