Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu

Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu

徒有其名的多语言能力?大语言模型在乌尔都语中的文化与语言缺陷

Abstract: Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the generation of multilingual LLMs when used for the task of story generation. We consider Urdu language as a representative low-resource language.

摘要: 多语言大语言模型(LLMs)正越来越多地被用于开放式文本生成,然而它们在低资源语言中的表现仍未得到充分研究。在这项工作中,我们探讨了多语言大语言模型在执行故事生成任务时,其生成内容的准确性和可靠性如何。我们以乌尔都语作为低资源语言的代表进行了研究。

We generate Urdu-Stories, a corpus of 93 stories generated using three contemporary LLMs (GPT-5.1, Qwen-3-Max, DeepSeek-3.1). We manually annotate the errors present in them under a nine-label linguistic, semantic, and cultural taxonomy. Our notable findings suggest that LLMs often make basic errors of grammar and semantics. The stories lack coherence, have unnatural repetition and show pervasive cultural shallowness.

我们构建了“Urdu-Stories”语料库,其中包含使用三种当代大语言模型(GPT-5.1、Qwen-3-Max、DeepSeek-3.1)生成的 93 个故事。我们根据包含语言、语义和文化维度的九类标签,对这些故事中存在的错误进行了人工标注。我们的重要发现表明,大语言模型经常犯基础的语法和语义错误。这些故事缺乏连贯性,存在不自然的重复,并表现出普遍的文化浅薄。

We further show using few-shot prompting that the cultural and context errors largely remain unresolved. Our findings highlight the limitations of current LLMs as a reliable source of content generation and information retrieval for low-resource languages.

我们进一步通过少样本提示(few-shot prompting)证明,这些文化和语境错误在很大程度上仍未得到解决。我们的研究结果凸显了当前大语言模型在作为低资源语言的内容生成和信息检索可靠来源方面的局限性。