Which India Survives Translation? Narrative Homogenisation Across Indian Oral Traditions in LLMs
Which India Survives Translation? Narrative Homogenisation Across Indian Oral Traditions in LLMs
翻译之后,哪个印度留存了下来?大语言模型中印度口述传统的叙事同质化研究
Abstract: Large language models (LLMs) are trained predominantly on English-language internet text that over-represents certain cultural narratives, raising concerns that models flatten the diversity of non-Western storytelling traditions into a single homogenized archetype.
摘要: 大语言模型(LLM)主要基于英语互联网文本进行训练,这些文本过度代表了特定的文化叙事,引发了人们的担忧:模型可能会将非西方叙事传统的多样性抹平,将其简化为单一的同质化原型。
We present a pilot computational study examining this across three maximally distinct Indian regional oral and literary traditions: the Rajasthani Pabuji epic, classical Tamil Sangam poetry, and Bengali folk tales. We collected authentic reference corpora for each tradition (11, 21, and 10 passages respectively) and prompted two LLMs (Claude Sonnet and Gemini) with 54 generation requests spanning three prompt types per tradition - generic, culturally specific, and regional-language.
我们开展了一项初步计算研究,考察了三种差异极大的印度区域口述及文学传统:拉贾斯坦邦的帕布吉(Pabuji)史诗、古典泰米尔桑伽姆(Sangam)诗歌以及孟加拉民间故事。我们为每种传统收集了真实的参考语料库(分别为 11、21 和 10 段文本),并向两个大模型(Claude Sonnet 和 Gemini)发送了 54 个生成请求,涵盖了每种传统下的三种提示类型——通用型、文化特定型和区域语言型。
Using Sentence-BERT embeddings and cosine similarity, we measure reference drift (how closely outputs track their own tradition’s authentic texts relative to the other two) and cross-tradition convergence (how similar outputs are across traditions). We find that while outputs remain closer to their own tradition’s reference than to others, cross-tradition similarity is high (0.52-0.66) relative to what the traditions’ genuine distance would predict, indicating partial homogenisation.
利用 Sentence-BERT 嵌入和余弦相似度,我们测量了“参考漂移”(输出内容相对于其他两种传统,在多大程度上贴合其自身传统的真实文本)以及“跨传统趋同性”(不同传统间的输出相似度)。研究发现,尽管输出内容在自身传统参考文本上的表现仍优于其他传统,但跨传统的相似度较高(0.52-0.66),这超出了基于这些传统真实差异的预期,表明存在部分同质化现象。
Unexpectedly, prompting in the regional language (Hindi, Tamil, or Bengali) consistently reduced fidelity to the authentic tradition relative to English prompting, by as much as 27 percentage points for Rajasthani and Bengali traditions. We discuss this against conflicting prior results on multilingual prompting and argue it reflects a difference between eliciting general cultural diversity and simulating one narrow, lesser-documented oral tradition.
出乎意料的是,与英语提示相比,使用区域语言(印地语、泰米尔语或孟加拉语)进行提示反而持续降低了对真实传统的忠实度,在拉贾斯坦和孟加拉传统中,这一降幅高达 27 个百分点。我们结合此前关于多语言提示的矛盾研究结果讨论了这一现象,并认为这反映了“激发一般文化多样性”与“模拟单一狭窄且记录较少的口述传统”之间的差异。
We position this pilot as a lightweight, scalable complement to recent large-scale human-annotation studies of Indian cultural misrepresentation in LLM-generated stories, as part of a broader doctoral research program.
我们将这项初步研究定位为对近期大规模人工标注研究(关于 LLM 生成故事中印度文化失真问题)的一种轻量级、可扩展的补充,这也是更广泛博士研究项目的一部分。