A Computational Approach to Measuring Semantic Change in Sanskrit Literature
A Computational Approach to Measuring Semantic Change in Sanskrit Literature
梵文文献语义演变测量的计算方法
Abstract: Diachronic word embeddings have become the modern standard for tracking semantic change, yet they have been largely validated on modern, high-resource, and well-segmented languages. This paper tests whether the paradigm transfers to Sanskrit, an ancient, low-resource language whose phonological fusion (sandhi), morphological inflection, compounding, and polysemy pose a unique challenge.
摘要: 历时性词嵌入(Diachronic word embeddings)已成为追踪语义演变的现代标准,但其主要是在现代、高资源且分词良好的语言上进行验证的。本文旨在测试这一范式是否适用于梵文——一种古老的低资源语言。梵文的语音连读(sandhi)、形态屈折、复合词结构以及多义性,为其语义分析带来了独特的挑战。
I assemble a 2.7M-token corpus spanning four canonical periods, recover word boundaries with a neural byte-level sandhi splitter and lemmatizer, and train per-period embeddings across configurations. To evaluate the system, I curate a validation set from historical scholarship and test recovery directionally with anchor displacement.
我构建了一个涵盖四个经典时期、包含 270 万个词元的语料库,利用神经字节级连读切分器(sandhi splitter)和词形还原器恢复了词边界,并针对不同配置训练了各时期的词嵌入模型。为了评估该系统,我从历史文献研究中整理了一套验证集,并通过锚点位移(anchor displacement)对语义恢复的方向性进行了测试。
Of 21 testable shifts, 19 move in the philologically attested direction (sign test, p=0.00011). I further show which configuration the language forces and comment on opportunities for improvement.
在 21 个可测试的语义偏移中,有 19 个与文献学证实的演变方向一致(符号检验,p=0.00011)。此外,我还展示了该语言特性所要求的特定配置,并对未来的改进方向提出了见解。