Training-Time Explainability for Multilingual Hate Speech Detection: Aligning Model Reasoning with Human Rationales
Training-Time Explainability for Multilingual Hate Speech Detection: Aligning Model Reasoning with Human Rationales
多语言仇恨言论检测的训练时可解释性:将模型推理与人类逻辑对齐
Abstract: Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation. Such systems, though accurate, remain opaque and risk bias, over-censorship, or under-moderation, particularly when detached from sociocultural context. 摘要: 针对穆斯林群体的网络仇恨往往以文化编码的多语言形式出现,从而逃避传统的人工智能审核。此类系统虽然准确,但仍然是不透明的,并存在偏见、过度审查或审查不足的风险,特别是在脱离社会文化背景的情况下。
We propose a training-time explainability framework that aligns model reasoning with human-annotated rationales, improving both classification performance and interpretability. Our approach is evaluated on HateXplain (English) and BullySent (Hinglish), reflecting the prevalence of anti-Muslim hate across both languages. 我们提出了一种训练时的可解释性框架,将模型推理与人类标注的逻辑(rationales)对齐,从而同时提升分类性能和可解释性。我们的方法在 HateXplain(英语)和 BullySent(印地语-英语混合语)数据集上进行了评估,反映了反穆斯林仇恨在这两种语言中的普遍性。
Using LIME, Integrated Gradients, Grad X Input, and attention, we assess accuracy, explanation quality, and cross-method agreement. Results show that gradient- and attention-based regularization improve F-scores, enhance plausibility and faithfulness, and capture culturally specific cues for detecting implicit anti-Muslim hate, offering a path toward multilingual, culturally aware content moderation. 通过使用 LIME、集成梯度(Integrated Gradients)、Grad X Input 和注意力机制,我们评估了准确性、解释质量以及跨方法的一致性。结果表明,基于梯度和注意力的正则化提高了 F-分数,增强了合理性和忠实度,并捕捉到了用于检测隐性反穆斯林仇恨的文化特定线索,为实现多语言、具备文化意识的内容审核提供了一条路径。