VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages

VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages

VakyArth:评估大语言模型在印度语言中的语用能力

Abstract: Real-world communication often requires pragmatic reasoning: interpreting meanings implied through context and cultural convention rather than stated literally. Existing pragmatic evaluation remains largely limited to English and high-resource languages, leaving Indic languages unexplored despite their linguistic and cultural diversity. 摘要: 现实世界的交流通常需要语用推理能力:即解读通过语境和文化习俗隐含的意义,而非仅仅理解字面意思。现有的语用评估主要局限于英语和高资源语言,尽管印度语言具有丰富的语言和文化多样性,但相关研究仍处于空白。

We introduce VakyArth, the first pragmatic benchmark for Indic languages, designed as a diagnostic evaluation covering Hindi, Punjabi, Tamil, and Malayalam. VakyArth evaluates models across five phenomena: deixis, speech acts, implicature, social pragmatics, and coherence; through multiple-choice questions, natural language inference, and translation, with all items authored by native speakers. 我们推出了 VakyArth,这是首个针对印度语言的语用基准测试,旨在对印地语、旁遮普语、泰米尔语和马拉雅拉姆语进行诊断性评估。VakyArth 通过多项选择题、自然语言推理和翻译等方式,从指示词、言语行为、隐含意义、社会语用和连贯性这五个维度对模型进行评估,所有测试项目均由母语人士编写。

Across multilingual large language models (LLMs) of varying families and sizes, we find consistent failures on pragmatic meanings rooted in Indic linguistic and cultural conventions. Our analysis shows systematic differences across languages and tasks: MCQ accuracy exceeds NLI accuracy in all model-language combinations, translation performance does not reliably track pragmatic understanding, and Indo-Aryan languages show a translation advantage over Dravidian languages. 在不同系列和规模的多语言大语言模型(LLM)中,我们发现模型在处理植根于印度语言和文化习俗的语用意义时,表现出持续的失败。我们的分析显示,不同语言和任务之间存在系统性差异:在所有模型与语言的组合中,多项选择题(MCQ)的准确率均高于自然语言推理(NLI)的准确率;翻译表现并不能可靠地反映语用理解能力;此外,印度-雅利安语系在翻译任务上表现出优于达罗毗荼语系的优势。

We further show that automatic translation metrics can miss fluent but pragmatically unfaithful outputs, especially for implicature and deixis. 我们进一步指出,自动翻译指标可能会忽略那些虽然流畅但语用不准确的输出,特别是在处理隐含意义和指示词时。