Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms

Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms

被识别却无法生成:文化特定亲属称谓的生成基准研究

Abstract: Current literature evaluates large language models (LLMs) on multilingual kinship understanding using multiple choice benchmarks, treating it as a recognition problem. We instead prompt five open-weight LLMs to generate kinship terms in three non-Western languages (Hindi, Tamil, and Korean) across two communicative tasks and pair this with a matched option-supported selection baseline. 摘要: 现有的文献多通过多项选择基准来评估大型语言模型(LLM)对多语言亲属关系的理解,将其视为一种识别问题。与之不同,我们要求五个开源权重的大型语言模型在两项交流任务中,生成三种非西方语言(印地语、泰米尔语和韩语)的亲属称谓,并将其与匹配的选项支持选择基准进行对比。

On identical relation language cells, GPT-OSS-120B selects the correct term in 90.67% of 75 valid cells but produces an accepted term in 36.00% of the corresponding attempts; Llama-3.3-70B shows the same pattern (77.92% versus 24.24%). Since the four-option condition displays the candidate terms and does not require script production, the difference is interpreted as an evaluation format gap rather than direct proof that lexical knowledge is intact. 在相同的关系语言单元格中,GPT-OSS-120B 在 75 个有效单元格中选出正确术语的准确率为 90.67%,但在相应的生成尝试中,仅有 36.00% 的尝试产生了可接受的术语;Llama-3.3-70B 也表现出同样的模式(77.92% 对比 24.24%)。由于四选项条件会显示候选术语且不需要书写脚本,这种差异被解释为评估格式上的差距,而非词汇知识完好的直接证据。

On explicitly specified L3 prompts, accuracy varies sharply, from GLM-5.1 at 72.29% to Llama-3.3-70B at 24.24%. The paternal-lineage advantage is language-specific; it is large in Hindi but weak or reversed in Korean, while Tamil shared-term pairs provide a control for measurement variation. 在明确指定的 L3 提示下,准确率差异巨大,从 GLM-5.1 的 72.29% 到 Llama-3.3-70B 的 24.24% 不等。父系血缘优势具有语言特异性:在印地语中表现显著,但在韩语中较弱或呈现相反趋势,而泰米尔语的共享术语对则为测量差异提供了对照。

These results show that culturally specific kinship generation remains difficult even when the relationship is explicitly stated and motivate generation-based evaluation alongside multiple-choice testing. 这些结果表明,即使明确说明了亲属关系,文化特定的亲属称谓生成仍然具有挑战性,这促使我们在多项选择测试之外,进一步开展基于生成的评估研究。