A large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United States

A large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United States

美国宗教广播网络流媒体转录语料库

Abstract: Religious radio is a widespread but understudied form of mass communication in the United States, and content-level analysis of it has been constrained by the absence of large-scale transcript data. 摘要: 宗教广播是美国一种广泛存在但研究不足的大众传播形式,由于缺乏大规模的转录数据,对其进行内容层面的分析一直受到限制。

This Data Descriptor presents a corpus of transcribed English-language religious radio broadcasts captured from live webstreams over a one-month period in July 2025. 本数据描述符展示了一个英语宗教广播转录语料库,该语料库采集自 2025 年 7 月为期一个月内的实时网络流媒体。

Fifteen-minute segments were recorded on a rolling schedule from 785 distinct streams, which together rebroadcast the signals of more than two thousand AM and FM stations, yielding over 700,000 recordings and more than 60 million diarized lines of speech. 研究人员通过滚动调度,从 785 个不同的流媒体中录制了 15 分钟一段的音频片段。这些流媒体共同转播了超过两千个 AM 和 FM 电台的信号,最终产生了超过 70 万条录音和超过 6000 万行经过说话人日志(diarized)处理的语音文本。

Each recording was transcribed and speaker-diarized with an automated pipeline, and segmented and labeled by programming format and topic using a large language model. 每条录音都通过自动化流程进行了转录和说话人日志处理,并利用大语言模型按节目形式和主题进行了分段与标注。

The corpus is organized as linked tables of stream metadata, recording metadata, and transcript lines. 该语料库以流媒体元数据、录音元数据和转录文本行的关联表形式进行组织。

It supports descriptive study of religious broadcasting across regions and traditions, analysis of how social and political issues are discussed in religious media, and speech-processing research in an underrepresented domain. 它支持对不同地区和传统的宗教广播进行描述性研究,分析宗教媒体如何讨论社会和政治议题,并为语音处理研究提供了一个代表性不足领域的数据支持。