Confidence-Controlled XAI Auditing for Pedestrian Detection under Domain Shift
Confidence-Controlled XAI Auditing for Pedestrian Detection under Domain Shift
领域偏移下行人检测的置信度受控可解释人工智能(XAI)审计
Abstract: Explainability is increasingly required for perception models in intelligent vehicles, yet whether explanations remain faithful under driving domain shift is still poorly understood. This work audits post-hoc explanations of a fixed YOLOv8s pedestrian detector across PIE and JAAD using ROI-based D-Deletion, frozen confidence terciles, rank-based tests, bootstrap intervals, and Holm correction.
摘要: 智能车辆的感知模型对可解释性的要求日益提高,然而在驾驶领域发生偏移时,这些解释是否依然保持忠实度,目前尚不明确。本研究通过基于 ROI 的 D-Deletion(删除法)、固定置信度三分位数、基于秩的检验、自助法区间估计以及 Holm 校正,对跨 PIE 和 JAAD 数据集的固定 YOLOv8s 行人检测器的事后解释进行了审计。
The audit shows that deletion-based faithfulness is strongly coupled to detection strength at explanation time, with Spearman correlations between 0.70 and 0.82 for D-RISE, making naive confidence-stratified comparisons unreliable. After controlling for detection strength within fixed f0 bins, D-RISE faithfulness remains domain-dependent in the central f0 range, with PIE showing higher D-Deletion than JAAD and Holm-adjusted significance.
审计结果表明,基于删除法的忠实度与解释时的检测强度密切相关,D-RISE 的 Spearman 相关系数在 0.70 到 0.82 之间,这使得简单的置信度分层比较变得不可靠。在固定 f0 区间内控制检测强度后,D-RISE 的忠实度在中心 f0 范围内依然表现出领域依赖性,其中 PIE 的 D-Deletion 指标高于 JAAD,且具有 Holm 校正后的显著性。
A non-perturbative EigenCAM baseline is less faithful than D-RISE but also exhibits score coupling, suggesting that the effect is not specific to D-RISE and is related to the deletion-based evaluation setup. These results motivate confidence-controlled XAI audits for safety-critical perception under domain shift.
一种非扰动式的 EigenCAM 基准模型虽然比 D-RISE 的忠实度更低,但也表现出了分数耦合现象,这表明该效应并非 D-RISE 所特有,而是与基于删除法的评估设置有关。这些结果强调了在领域偏移下,针对安全关键型感知任务进行置信度受控的 XAI 审计的必要性。