Measuring the Cross-Lingual Comprehension Gap: How the language of the evidence shapes what language models understand

Measuring the Cross-Lingual Comprehension Gap: How the language of the evidence shapes what language models understand

衡量跨语言理解差距:证据语言如何塑造语言模型的理解能力

Abstract: Language models are often evaluated as though capabilities demonstrated in English remain equally available when the same content is presented in other languages. Traditional multilingual benchmarks rarely isolate language while holding content, question, reference answer, model, and evaluation unit constant.

摘要: 语言模型在评估时,往往被假设其在英语中展现的能力,在处理其他语言的相同内容时依然保持不变。传统的跨语言基准测试很少能在保持内容、问题、参考答案、模型和评估单元不变的情况下,单纯地隔离语言变量。

We define the Cross-Lingual Comprehension Gap (CLCG) as the reduction in response quality when the same content and question are presented in a target language rather than in English. Using ParallelQA-18, a professionally human-translated parallel corpus, we evaluate five models from five laboratories on a stratified sample of 150 articles across 18 languages (English reference; Portuguese high-resource baseline; 16 targets spanning Joshi et al. 2020 classes 0-4).

我们将“跨语言理解差距”(Cross-Lingual Comprehension Gap, CLCG)定义为:当相同的内容和问题以目标语言而非英语呈现时,模型响应质量的下降程度。我们使用由专业人工翻译的平行语料库 ParallelQA-18,对来自五个实验室的五个模型进行了评估,样本涵盖了 18 种语言的 150 篇文章(以英语为参考,葡萄牙语为高资源基准,其余 16 种目标语言涵盖了 Joshi 等人 2020 年定义的 0-4 类资源水平)。

A within-item design varies only passage language. The primary estimator contrasts English versus pooled target-language Token-F1 micro-means on higher-complexity open-ended questions, with article-cluster bootstrap intervals. The primary pooled CLCG is 0.078 (95% CI 0.072-0.084), about a 17% reduction relative to the English score; the equal-language macro summary is 0.077. Net of Portuguese, the macro gap is 0.016 (95% CI 0.013-0.020).

研究采用项内设计(within-item design),仅改变文章语言。主要评估指标对比了英语与目标语言在处理高复杂度开放式问题时的 Token-F1 微观平均值,并计算了文章聚类的自助法置信区间。主要的汇总 CLCG 为 0.078(95% 置信区间 0.072-0.084),相对于英语得分下降了约 17%;等权语言宏观汇总值为 0.077。剔除葡萄牙语后,宏观差距为 0.016(95% 置信区间 0.013-0.020)。

Language-level CLCG is negatively associated with Joshi resource class (rho = -0.594, p = 0.015, n = 16). In blinded paired human evaluations, higher-resource responses are preferred in 61.6% of decisive judgments (estimated preference probability 0.655, 95% CI 0.558-0.741). Capabilities shown in English should not be assumed to transfer equally to other languages; English-centered evaluations may overestimate quality for users of low-resource languages.

语言层面的 CLCG 与 Joshi 资源等级呈负相关(rho = -0.594, p = 0.015, n = 16)。在盲测配对人工评估中,61.6% 的决定性判断更倾向于高资源语言的回答(估计偏好概率为 0.655,95% 置信区间 0.558-0.741)。我们不应假设在英语中展现的能力能同等地迁移到其他语言;以英语为中心的评估可能会高估模型对低资源语言用户的服务质量。