Enabling Vision and Cross-Modal Learning for Multimodal Stroke Recurrence Prediction: An Interpretable Two-Step Framework

Enabling Vision and Cross-Modal Learning for Multimodal Stroke Recurrence Prediction: An Interpretable Two-Step Framework

赋能视觉与跨模态学习:用于多模态中风复发预测的可解释两步框架

Abstract: Multimodal stroke recurrence prediction requires effective integration of heterogeneous clinical and imaging data, yet modality imbalance often causes models to over-rely on dominant modalities and underutilize complementary information. 摘要: 多模态中风复发预测需要有效整合异构的临床和影像数据,然而模态不平衡往往导致模型过度依赖主导模态,而未能充分利用互补信息。

While self-supervised pretraining and selective parameter freezing are commonly employed to improve representation learning and fine-tuning stability, their effect on modality contributions and cross-modal behavior in multimodal medical models remains largely unexplored. 虽然自监督预训练和选择性参数冻结常被用于改善表征学习和微调稳定性,但它们对多模态医学模型中模态贡献和跨模态行为的影响在很大程度上仍未得到探索。

In this work, we investigate whether image pretraining on 3D CTA scans reduces modality imbalance and improves cross-modal integration for stroke recurrence prediction, a clinically critical task we recently addressed. 在这项工作中,我们研究了在 3D CTA 扫描图像上进行预训练是否能减少模态不平衡,并改善中风复发预测(我们近期解决的一项临床关键任务)中的跨模态整合效果。

To this end, two multimodal neural networks are pretrained in a self-supervised manner and subsequently fine-tuned using two distinct freezing strategies. Their performance and modality utilization are compared against both the baseline model from our previous work and models trained entirely from scratch in this study. 为此,我们以自监督方式预训练了两个多模态神经网络,并随后使用两种不同的冻结策略进行微调。我们将它们的性能和模态利用率与我们之前工作中的基准模型以及本研究中完全从头训练的模型进行了对比。

Our results demonstrate that self-supervised pretraining enables more effective utilization of the multimodal image-tabular dataset, outperforming both the prior baseline and all non-pretrained models. Notably, the best-performing Vision Transformer based neural network successfully overcomes unimodal collapse. 结果表明,自监督预训练能够更有效地利用多模态图像-表格数据集,其表现优于之前的基准模型和所有未经预训练的模型。值得注意的是,表现最好的基于 Vision Transformer 的神经网络成功克服了单模态坍塌问题。

Synergy analysis reveals significant interactions between vision and both gender and CHD, suggesting clinically relevant patterns for stroke recurrence. Overall, our findings demonstrate that self-supervised pretraining and strategic fine-tuning support more balanced modality utilization and enable meaningful cross-modal interactions. Code is publicly available at this https URL. 协同分析揭示了视觉与性别及冠心病(CHD)之间存在显著的交互作用,这表明了中风复发中具有临床意义的模式。总的来说,我们的研究结果证明,自监督预训练和策略性微调支持了更平衡的模态利用,并实现了有意义的跨模态交互。代码已在链接中公开。