There is No Theoretical Curse of Multilinguality For Embedding Space Structure

There is No Theoretical Curse of Multilinguality For Embedding Space Structure

嵌入空间结构不存在理论上的“多语言诅咒”

Abstract: A central goal of multilingual NLP is to achieve high monolingual performance per language and cross-lingual alignment for large-scale language coverage with a multilingual model. The curse of multilinguality describes the phenomenon of degradation in multilingual model performance as we increase language coverage, posing a threat to the above goal.

摘要: 多语言自然语言处理(NLP)的一个核心目标是,利用多语言模型实现每种语言的高单语性能,并实现大规模语言覆盖下的跨语言对齐。“多语言诅咒”(curse of multilinguality)描述了随着语言覆盖范围的扩大,多语言模型性能出现下降的现象,这对上述目标构成了威胁。

This paper asks whether multilingual embedding spaces are inherently incapable of achieving perfect multilinguality without a prohibitive increase in required capacity. We first formalize the goal of “perfect multilinguality”, embodied in two multilinguality conditions.

本文探讨了多语言嵌入空间是否在本质上无法在不大幅增加所需容量的情况下实现“完美多语言”。我们首先将“完美多语言”的目标形式化,并将其具体化为两个多语言条件。

We then prove that the minimum dimensionality required for perfect multilinguality grows only logarithmically in the number of languages. That is, we show that there is no theoretical curse of multilinguality for embedding space structure. This suggests that the empirical curse of multilinguality is a result of real world data and training conditions.

随后我们证明,实现完美多语言所需的最小维度仅随语言数量呈对数增长。也就是说,我们证明了在嵌入空间结构方面,并不存在理论上的“多语言诅咒”。这表明,经验观察到的“多语言诅咒”实际上是现实世界数据和训练条件所导致的结果。

We back this understanding with a small-scale empirical study. Our paper provides the first theoretical and intrinsic perspective on the curse of multilinguality, with implications for the scientific understanding of this phenomenon.

我们通过一项小规模的实证研究支持了这一理解。本文首次从理论和内在视角对“多语言诅咒”进行了探讨,并对科学界理解这一现象具有重要意义。