TutlAit v1: a crowdsourced Moroccan Tamazight speech dataset with Arabic transcriptions and regional accent labels

TutlAit v1: a crowdsourced Moroccan Tamazight speech dataset with Arabic transcriptions and regional accent labels

TutlAit v1:一个带有阿拉伯语转录和区域口音标签的众包摩洛哥塔马齐格语语音数据集


Abstract: Tamazight (Amazigh) is, together with Arabic, one of the two official languages of Morocco, yet it remains severely under-resourced for speech technology: publicly available labelled audio is scarce, generally lacks information on the regional variety spoken, and is often of uneven transcription quality.

摘要: 塔马齐格语(Amazigh)与阿拉伯语同为摩洛哥的两种官方语言,但在语音技术领域,它仍然极度缺乏资源:公开可用的标注音频稀缺,通常缺乏所讲区域变体的信息,且转录质量往往参差不齐。


This article describes the TutlAit dataset, a corpus of Moroccan Tamazight speech paired with Modern Standard Arabic text and explicit regional accent labels. The data were collected with TutlAit, a purpose-built crowdsourcing web application (React 18 front end, Django 5 / Django REST Framework back-end, PostgreSQL database).

本文介绍了 TutlAit 数据集,这是一个包含摩洛哥塔马齐格语语音、现代标准阿拉伯语文本以及明确区域口音标签的语料库。这些数据是通过专门构建的众包 Web 应用程序 TutlAit(前端采用 React 18,后端采用 Django 5 / Django REST Framework,数据库采用 PostgreSQL)收集的。


Native speakers recruited through targeted LinkedIn and Instagram campaigns created an account, declared their regional variety (Atlas, Souss, Rif or other) and demographic information, and then contributed through two workflows: Text-to-Audio, in which an Arabic sentence is displayed and the volunteer records its oral Tamazight rendering in the browser, and Audio-to-Text, in which a Tamazight excerpt is played and the volunteer types its Arabic transcription.

通过 LinkedIn 和 Instagram 定向活动招募的母语人士在创建账户后,需声明其所属的区域变体(阿特拉斯、苏斯、里夫或其他)及人口统计信息,随后通过两种工作流进行贡献:一是“文本转语音”,即显示一句阿拉伯语,志愿者在浏览器中录制其塔马齐格语口述版本;二是“语音转文本”,即播放一段塔马齐格语片段,志愿者输入其阿拉伯语转录内容。


A complementary set of segments was obtained from freely accessible Tamazight audiovisual media, segmented and annotated with ELAN and imported through a bulk CSV/ZIP pipeline. Every upload is converted server-side to 16kHz mono WAV, hashed with SHA-256 for duplicate rejection, checked for duration bounds and validated by an administrator.

另一部分补充片段取自可自由访问的塔马齐格语视听媒体,通过 ELAN 进行分段和标注,并经由批量 CSV/ZIP 管道导入。每次上传的文件都会在服务器端转换为 16kHz 单声道 WAV 格式,通过 SHA-256 进行哈希处理以剔除重复项,并经过时长限制检查及管理员验证。


The dataset contains 13,384 audio files totalling 75,231 seconds (approximately 20.9 hours, about 3.01GB). The Atlas variety accounts for 9,956 files (14.08h) and the Souss variety for 3,378 files (6.75h); small Rif (22 files) and Kabyle (28 files) subsets are also included. The corpus can be reused for speech recognition, speech translation and accent identification for Moroccan Tamazight.

该数据集包含 13,384 个音频文件,总时长为 75,231 秒(约 20.9 小时,约 3.01GB)。其中阿特拉斯变体占 9,956 个文件(14.08 小时),苏斯变体占 3,378 个文件(6.75 小时);此外还包括少量的里夫语(22 个文件)和卡拜尔语(28 个文件)子集。该语料库可用于摩洛哥塔马齐格语的语音识别、语音翻译和口音识别研究。