BharatGather: A Culturally-Informed Benchmark Dataset for Misinformation and Fake News Detection in Indian Public Events
BharatGather: A Culturally-Informed Benchmark Dataset for Misinformation and Fake News Detection in Indian Public Events
BharatGather:一个用于印度公共事件虚假信息和假新闻检测的文化感知基准数据集
Abstract: Large-scale public events, such as religious festivals, political rallies, and cultural gatherings, are increasingly vulnerable to the rapid dissemination of misinformation, posing substantial risks to public safety and social cohesion. 摘要: 宗教节日、政治集会和文化聚会等大型公共活动越来越容易受到虚假信息快速传播的影响,这对公共安全和社会凝聚力构成了重大风险。
While automated fake news detection has seen significant methodological progress, existing benchmarks frequently fail to capture the socio-cultural nuances and event-specific dynamics characteristic of the Indian context. 尽管自动化假新闻检测在方法论上已取得显著进展,但现有的基准测试往往无法捕捉印度背景下特有的社会文化细微差别和事件动态。
This paper introduces BharatGather, a curated, multi-source dataset specifically engineered for binary misinformation classification within the ecosystem of Indian mass gatherings. 本文介绍了 BharatGather,这是一个经过精心策划的多源数据集,专门为印度大型集会生态系统内的二元虚假信息分类而设计。
The corpus comprises 14,646 records constructed through a hybrid pipeline involving systematic web scraping of prominent fact-checking platforms, multimedia transcript extraction, and Large Language Model (LLM)-mediated synthetic augmentation to ensure narrative diversity. 该语料库包含 14,646 条记录,通过混合流程构建而成,该流程涉及对知名事实核查平台的系统性网络抓取、多媒体转录提取,以及由大语言模型(LLM)介导的合成增强,以确保叙事的多样性。
By providing a resource tailored to the unique complexities of event-aware misinformation in India, this work facilitates the development of culturally informed detection systems and establishes a rigorous benchmark for evaluating their performance in high-stakes public environments. 通过提供一种针对印度事件感知虚假信息独特复杂性而量身定制的资源,本研究促进了文化感知检测系统的开发,并为评估其在高风险公共环境中的性能建立了严格的基准。
Authors: Parth Bramhecha, Smit Deshmukh, Sairaj Bodhale, Adwait Borate, Raviraj Joshi 作者: Parth Bramhecha, Smit Deshmukh, Sairaj Bodhale, Adwait Borate, Raviraj Joshi
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG) 学科: 计算与语言 (cs.CL);机器学习 (cs.LG)
arXiv ID: 2609.02895 arXiv 编号: 2609.02895