Framing by Wording, Framing by Selection: A Large-Scale Two-Dimensional Audit of French News Headlines, 2022-2025

Framing by Wording, Framing by Selection: A Large-Scale Two-Dimensional Audit of French News Headlines, 2022-2025

措辞框架与选题框架:2022-2025年法国新闻标题的大规模二维审计

Abstract: News headlines frame public issues both by what they select and by how they word it, yet computational framing work typically collapses these operations into a single score. We introduce a two-dimensional framework that separates salience framing, measured through four wording devices (loaded vocabulary, blame attribution, threat framing, rhetorical question), from selection framing, measured through outlet-level story-form and high-charge distributions.

摘要: 新闻标题通过“选择什么内容”以及“如何措辞”来构建公共议题,然而目前的计算框架研究通常将这两种操作合并为一个单一的分数。我们引入了一个二维框架,将“显著性框架”(salience framing)与“选题框架”(selection framing)分离开来。前者通过四种措辞手段(加载词汇、归咎责任、威胁框架、反问句)进行衡量,后者则通过媒体层面的报道形式和高强度分布进行衡量。

We build a 10,000-headline French supervision set using three LLM annotators with majority-vote resolution and human arbitration, validate the labels against two annotator-independent blind human studies, and apply the strongest classifier to 902,111 deduplicated headlines from 25 French outlets (2022-2025).

我们利用三个大语言模型(LLM)标注器构建了一个包含10,000个法语标题的监督数据集,并采用了多数投票决议和人工仲裁机制。我们通过两项独立于标注者的人工盲测验证了标签的准确性,并将表现最优的分类器应用于来自25家法国媒体的902,111条去重标题(2022-2025年)。

Three main findings emerge. First, salience and selection divergence are positively correlated yet leave nearly half of outlet-level variance unexplained, populating interpretively distinct off-diagonal cells in a four-cell outlet typology. Second, default classification thresholds systematically inflate corpus-level salience estimates; a precision-floor recalibration protocol corrects this distortion.

研究得出三个主要发现。首先,显著性和选题的差异呈正相关,但仍有近一半的媒体层面方差无法解释,这在四格媒体类型学中形成了具有不同解释意义的非对角线单元。其次,默认的分类阈值会系统性地夸大语料库层面的显著性估计;我们通过一种“精度下限校准协议”修正了这一偏差。

Third, group-mention analysis reveals sharply unequal salience contexts: headlines mentioning Jews, the Far-right, and Muslims carry the highest detected salience rates, which broad event-context composition does not fully explain (residuals are descriptive, not same-event causal estimates; per-group lexicon precision is reported alongside). To our knowledge, this is the largest framing-focused French headline audit to date; we release the supervision set, lexicons, and analysis code.

第三,群体提及分析揭示了显著性语境存在严重的不平等:提及犹太人、极右翼和穆斯林的标题具有最高的显著性检出率,而广泛的事件背景构成并不能完全解释这一点(残差仅为描述性统计,而非同一事件的因果估计;文中同时报告了各群体的词汇精度)。据我们所知,这是迄今为止规模最大的针对法语新闻标题的框架审计研究;我们已公开了监督数据集、词汇表及分析代码。