Wazobia Eval: A Benchmark for Nigerian Pidgin Emotion Understanding, Sarcasm Detection, and Cultural Reasoning
Wazobia Eval: A Benchmark for Nigerian Pidgin Emotion Understanding, Sarcasm Detection, and Cultural Reasoning
Wazobia Eval:针对尼日利亚皮钦语情感理解、讽刺检测及文化推理的基准测试
Abstract: Nigerian Pidgin is one of Africa’s most widely spoken languages, yet remains severely underrepresented in language model evaluation. Existing benchmarks primarily focus on translation, transcription, or generic sentiment analysis, leaving critical aspects of culturally grounded language understanding unmeasured.
摘要: 尼日利亚皮钦语(Nigerian Pidgin)是非洲使用最广泛的语言之一,但在语言模型评估中却严重缺乏代表性。现有的基准测试主要集中在翻译、转录或通用的情感分析上,导致对基于文化背景的语言理解这一关键方面的评估处于空白。
We introduce Wazobia Eval, a benchmark for evaluating Nigerian Pidgin emotion understanding, sarcasm detection, and cultural reasoning. The benchmark is built on a manually annotated dataset containing over 550 examples and a 16-category emotion taxonomy designed to capture culturally specific emotional registers that are not represented in conventional sentiment frameworks.
我们推出了 Wazobia Eval,这是一个用于评估尼日利亚皮钦语情感理解、讽刺检测和文化推理的基准测试。该基准建立在一个包含超过 550 个示例的手动标注数据集之上,并采用了一套 16 类情感分类法,旨在捕捉传统情感分析框架中未涵盖的、具有文化特异性的情感表达。
Wazobia Eval provides standardized evaluation protocols and benchmark tasks for assessing model performance on nuanced Nigerian language understanding. We present the benchmark design, annotation methodology, taxonomy development process, and preliminary pilot evaluation results. Our goal is to provide foundational evaluation infrastructure for Nigerian language AI and establish a reproducible benchmark for future research. The dataset is publicly available at this https URL.
Wazobia Eval 提供了标准化的评估协议和基准任务,用于评估模型在细微的尼日利亚语言理解方面的表现。我们介绍了该基准的设计、标注方法、分类法开发过程以及初步的试点评估结果。我们的目标是为尼日利亚语言人工智能提供基础的评估基础设施,并为未来的研究建立一个可复现的基准。该数据集已通过此链接公开。