SWORD: Wikidata-based Distortions Reveal Hidden Cross-Lingual Inconsistencies in LLM Factual Error Rejection

SWORD: Wikidata-based Distortions Reveal Hidden Cross-Lingual Inconsistencies in LLM Factual Error Rejection

SWORD:基于 Wikidata 的失真测试揭示了大语言模型在事实错误拒绝能力上的跨语言不一致性

Abstract: Modern LLMs demonstrate impressive multilingual performance, yet standard benchmarks primarily reward selecting correct answers rather than evaluating genuine factual understanding. 摘要: 现代大语言模型(LLM)展现出了令人印象深刻的多语言性能,但现有的标准基准测试主要奖励模型选择正确答案,而非评估其真实的事实理解能力。

We introduce Systematic Wikidata-based Object-Relation Distortion (SWORD), a benchmark that evaluates whether models consistently reject factual errors across languages. 我们引入了“基于 Wikidata 的系统性对象-关系失真”(Systematic Wikidata-based Object-Relation Distortion,简称 SWORD)基准测试,旨在评估模型在不同语言中是否能一致地拒绝事实错误。

SWORD generates syntactically well-formed but factually incorrect statements in eight widely spoken languages through controlled perturbations of Wikidata triples, ranging from random entity substitutions to semantically plausible property-based selections. SWORD 通过对 Wikidata 三元组进行受控扰动,生成了八种广泛使用语言中语法正确但事实错误的陈述,扰动方式涵盖了从随机实体替换到语义上看似合理的属性选择。

Our distortion-based evaluation surfaces two critical insights that remain entirely obscured by conventional benchmarks. 我们基于失真的评估揭示了两个关键见解,而这些见解在传统基准测试中完全被掩盖了。

First, models counterintuitively achieve higher accuracy on semantically plausible distortions than on nonsensical random substitutions, suggesting reliance on distributional familiarity rather than genuine factual verification. 首先,模型在语义上看似合理的失真陈述上,反而比在毫无意义的随机替换陈述上表现出更高的准确率,这表明模型更多是依赖分布上的熟悉度,而非真正的事实核查。

Second, models exhibiting comparable baseline accuracy across languages show substantial performance degradation specifically on (East) Asian languages when presented with distorted statements, with cross-lingual performance gaps reaching up to 28 percentage points (49% relative reduction) in some models. 其次,在面对失真陈述时,那些在基准测试中表现出跨语言准确性相当的模型,在(东)亚语言上却出现了显著的性能下降,部分模型的跨语言性能差距高达 28 个百分点(相对下降 49%)。

These findings demonstrate that multilingual factual reasoning involves asymmetric capabilities that aggregate accuracy metrics systematically obscure. 这些发现表明,多语言事实推理涉及非对称的能力差异,而汇总后的准确率指标系统性地掩盖了这些差异。