What Counts as a Mistake? Annotating Recitation Events in Quran Memorization Transcripts
What Counts as a Mistake? Annotating Recitation Events in Quran Memorization Transcripts
什么是错误?古兰经背诵转录本中诵读事件的标注研究
Abstract: Checking Quran recitation from an ASR transcript requires distinguishing unresolved mistakes from repetitions, repairs, opening formulas and accepted spelling differences. We report a completed human annotation of 100 production recording cases: 348 scored units and 162 localized events across ten combined labels. An executable evaluator scores labels and word positions together. A plain diff reaches label-aware F1 0.525 and localization F1 0.826; adapted production cleaner/alignment components reach 0.518 and 0.786, with exact-span F1 0.505 for both.
摘要: 通过自动语音识别(ASR)转录本核查古兰经诵读,需要将未解决的错误与重复、修正、开篇公式以及公认的拼写差异区分开来。我们报告了一项针对 100 个生产录音案例的人工标注成果:涵盖了 348 个评分单元和 162 个跨越十个组合标签的定位事件。一个可执行的评估器会对标签和单词位置进行联合评分。简单的差异对比(plain diff)在标签感知 F1 分数上达到 0.525,定位 F1 分数达到 0.826;经过调整的生产清理/对齐组件分别达到 0.518 和 0.786,两者的精确跨度(exact-span)F1 分数均为 0.505。
Correcting the adapter’s word coordinates recovers all five annotated repetition events, showing why annotation interfaces must be checked before interpreting baseline failures. In a preliminary pilot, eight single 20-minute runs across three coding agents and eight models span label-aware F1 0.143 to 0.892: seven land far above every baseline, and one collapses below the naive diff from a missing normalization step.
修正适配器的单词坐标后,成功恢复了所有五个已标注的重复事件,这表明在解读基线失败原因之前,必须先检查标注界面。在一项初步试点中,三名编码代理和八个模型进行了八次 20 分钟的单次运行,标签感知 F1 分数范围从 0.143 到 0.892:其中七次运行的结果远高于所有基线,而有一次因缺少归一化步骤而导致结果低于简单的差异对比。
Across the six, 970 of 972 gold-event instances draw an overlapping prediction, so what remains is not detection but convention: span extent, and the labels whose boundary is stipulated by adjudication rather than visible in the text. Seven of 162 events defeat all six same-day runs, five of them one orthographic rule, and the strongest run still misses the same ones. No run annotated before building, so the pilot measures the algorithm half of the task only.
在六次运行中,972 个黄金事件实例中有 970 个获得了重叠预测,因此剩下的问题不再是检测,而是惯例问题:即跨度范围,以及那些边界由裁定而非文本可见性所决定的标签。在 162 个事件中,有 7 个事件难倒了所有六次当日运行,其中 5 个涉及同一拼写规则,且表现最强的运行也未能识别出这些事件。由于在构建前没有任何运行进行过标注,因此该试点仅衡量了任务中算法那一半的表现。