AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
AHA-Memes:用于理解阿拉伯语模因中仇恨言论的细粒度多模态基准
Abstract: Hateful memes are a growing form of multimodal online harm, where hostile intent is often conveyed through the joint interpretation of images, text, cultural references, and implicit targets. While hateful meme detection has advanced in high-resource languages, Arabic remains underexplored, with existing meme resources focusing mainly on propaganda or coarse harmful-content labels.
摘要: 仇恨模因(Hateful memes)是一种日益严重的多模态网络危害形式,其敌意通常通过图像、文本、文化指涉和隐含目标的共同解读来传达。尽管仇恨模因检测在高资源语言中已取得进展,但阿拉伯语领域的研究仍显不足,现有的模因资源主要集中在宣传内容或粗略的有害内容标签上。
We introduce AHA-Memes (Arabic HAteful Memes), which is, to our knowledge, the first large-scale Arabic hateful meme benchmark with fine-grained, multi-label annotations. The dataset includes 5K manually annotated memes using a taxonomy that captures hate types, i.e., attack strategies. We further provide ~66K silver-labeled memes to support future studies.
我们推出了 AHA-Memes(阿拉伯语仇恨模因),据我们所知,这是首个具有细粒度、多标签注释的大规模阿拉伯语仇恨模因基准。该数据集包含 5,000 个经过人工注释的模因,并使用了一套涵盖仇恨类型(即攻击策略)的分类法。此外,我们还提供了约 6.6 万个银标签(自动标注)模因,以支持未来的研究。
We benchmark text-only, image-only, and late-fusion multimodal models, as well as few-shot in-context learning (ICL) and open- and closed-weight Vision-Language Models (VLMs) under zero-shot and fine-tuning settings. Our results establish strong baselines and highlight key challenges in culturally grounded Arabic hateful meme detection. We release the dataset, annotation guidelines, and evaluation scripts to support future research.
我们对纯文本、纯图像和后期融合多模态模型进行了基准测试,并在零样本(zero-shot)和微调设置下评估了少样本上下文学习(ICL)以及开源和闭源视觉语言模型(VLM)。我们的研究结果建立了强大的基准,并突显了在具有文化背景的阿拉伯语仇恨模因检测中所面临的关键挑战。我们已发布数据集、注释指南和评估脚本,以支持未来的研究。
WARNING: This paper contains examples that may be disturbing to readers.
警告: 本文包含可能令读者感到不安的示例。