How Much Does Corpus Choice Change Dependency-Distance Estimates?

How Much Does Corpus Choice Change Dependency-Distance Estimates?

语料库的选择对依存距离估计有多大影响?

Abstract: Dependency-distance estimates derived from a single corpus are routinely treated as properties of a language, yet this assumption has not been tested across independently compiled corpora.

摘要: 从单一语料库中得出的依存距离(Dependency-distance)估计值通常被视为语言的固有属性,然而这一假设尚未在独立编译的语料库之间得到验证。

We compared mean dependency-distance estimates across 38 same-language treebank pairs from Universal Dependencies v2.18, using concordance correlation, Bland-Altman analysis, and a twelve-specification multiverse design.

我们利用 Universal Dependencies v2.18 中的 38 对同语言树库(treebank),通过一致性相关分析(concordance correlation)、Bland-Altman 分析以及十二种规格的多重宇宙设计(multiverse design),比较了平均依存距离的估计值。

Cross-treebank agreement was moderate at best: substituting one treebank for another reversed nearly 40 percent of pairwise language orderings, and treebank choice accounted for roughly 29 percent of between-group variance.

不同树库之间的一致性充其量只能达到中等水平:将一个树库替换为另一个树库,会导致近 40% 的成对语言排序发生逆转,且树库的选择约占组间方差的 29%。

This disagreement substantially exceeded within-treebank sampling error and persisted across all twelve preprocessing specifications.

这种差异显著超过了树库内部的抽样误差,并且在所有十二种预处理规格中均持续存在。

Nevertheless, every treebank confirmed dependency-length minimization (normalized ratio below 1).

尽管如此,每个树库都证实了依存长度最小化(Dependency-length minimization, DLM)的存在(归一化比率低于 1)。

The data are more consistent with MDD as a corpus-conditioned composite of grammatical, register, and annotation factors than as a stable language-level parameter: the qualitative DLM universal survives corpus substitution, but the ordinal cross-linguistic ranking does not.

数据表明,MDD(平均依存距离)更像是一个受语料库条件影响的、由语法、语域和标注因素组成的复合体,而非一个稳定的语言级参数:定性的 DLM 普遍性在语料库替换后依然成立,但跨语言的序数排名则不然。