MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models

MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models

MemeCULT-1K:多模态模型南亚文化背景与幽默理解基准测试

Abstract: Meme understanding goes beyond recognizing visual content or literal text; it requires implicit cultural knowledge and pragmatic inference that most vision-language models still lack.

摘要: 对模因(Meme)的理解不仅仅是识别视觉内容或字面文本;它需要大多数视觉-语言模型目前仍缺乏的隐性文化知识和语用推理能力。

We introduce MemeCULT-1K, a multilingual benchmark of 1,000 South Asian memes in Bengali, English, and Hindi, where each meme is paired with a cultural context note and three human-written explanations, along with a supplementary set of 54 Bengali regional dialect memes.

我们推出了 MemeCULT-1K,这是一个包含 1,000 个南亚模因的多语言基准测试集,涵盖孟加拉语、英语和印地语。每个模因都配有文化背景注释和三条人工撰写的解释,此外还补充了 54 个孟加拉地区方言模因。

We evaluate thirteen popular Vision Language Models (VLMs) under two settings: meme-only and context-aware. Providing minimal cultural context yields consistent gains across all models and languages: mean SBERT similarity improves from 44.6 to 56.4 (+11.8), BLEURT from 37.3 to 42.3 (+5.0), and LLM-as-a-Judge scores from 2.57 to 3.43 out of 5 (+0.86).

我们在“仅模因”和“上下文感知”两种设置下评估了 13 种主流视觉语言模型(VLM)。提供最基本的文化背景后,所有模型和语言的表现均有持续提升:平均 SBERT 相似度从 44.6 提高到 56.4(+11.8),BLEURT 分数从 37.3 提高到 42.3(+5.0),“大模型作为裁判”(LLM-as-a-Judge)的评分从 2.57 提高到 3.43(满分 5 分,+0.86)。

Fine-grained error analysis reveals that closed-source models fail mainly on entity and reference misidentification, while open-source models are bottlenecked by broader cultural knowledge gaps, with linguistic and phonological failures proving the most context-resistant across both.

细粒度的错误分析显示,闭源模型的主要失败点在于实体和指代识别错误,而开源模型则受限于更广泛的文化知识缺口;其中,语言和语音层面的错误在两类模型中都表现出最强的“上下文抗性”(即难以通过提供上下文来纠正)。

These results highlight the difficulty of culturally grounded meme understanding and motivate future work on explicit cultural knowledge integration. Our dataset and code are publicly available at TawsifDipto17/MemeCULT-1K.

这些结果凸显了基于文化背景理解模因的难度,并为未来整合显性文化知识的研究提供了动力。我们的数据集和代码已在 TawsifDipto17/MemeCULT-1K 公开。