Complementary rPPG-Derived and Lip-Region Frequency Cues for Talking-Face Deepfake Detection
Complementary rPPG-Derived and Lip-Region Frequency Cues for Talking-Face Deepfake Detection
基于 rPPG 衍生信号与唇部频率特征互补的说话人脸深度伪造检测
Abstract: Talking-face (TF) deepfakes are detected unevenly by rPPG-based methods across generators. We study two lightweight visual-only cues, rPPG-derived waveforms extracted by RhythmFormer and lip-region discrete cosine transform (DCT) coefficients, on the seven TF methods of Celeb-DF++ under a subject-independent protocol.
摘要: 基于 rPPG(远程光电容积脉搏波描记法)的方法在检测不同生成器生成的说话人脸(TF)深度伪造时,效果参差不齐。我们研究了两种轻量级的纯视觉特征:由 RhythmFormer 提取的 rPPG 衍生波形,以及唇部区域的离散余弦变换(DCT)系数。我们在 Celeb-DF++ 数据集的七种 TF 方法上,采用受试者独立协议对这些特征进行了评估。
In-domain, lip-region DCT matches or exceeds the rPPG-derived 1D ResNet on every method except SadTalker, and Concat fusion reaches AUC 0.891 against 0.824 and 0.827 for the unimodal baselines. Under leave-one-generator-out evaluation the cues split: each transfers clearly better to three held-out methods, and IP-LAP is near chance for both. Concat averages 0.798 but falls below rPPG alone where DCT transfers poorly, so static fusion only partly exploits this complementarity.
在域内测试中,除 SadTalker 外,唇部区域 DCT 在所有方法上的表现均达到或超过了 rPPG 衍生的 1D ResNet;融合后的模型 AUC 达到 0.891,而单模态基准模型仅为 0.824 和 0.827。在“留一生成器”交叉验证评估中,两种特征表现出差异:每种特征在三个留出方法上的迁移效果明显更好,而两者在 IP-LAP 方法上的表现均接近随机水平。融合后的平均 AUC 为 0.798,但在 DCT 迁移效果较差的情况下,其表现低于单独使用 rPPG,这说明静态融合仅部分利用了这种互补性。
Lip-region DCT outperforms full-face DCT on six of seven methods. We treat the rPPG-derived signal as an empirical cue and do not claim it is cardiac in origin.
在七种方法中的六种里,唇部区域 DCT 的表现优于全脸 DCT。我们将 rPPG 衍生信号视为一种经验性特征,并不主张其具有心脏生理起源。