Holtercare-Bench: A Multimodal Benchmark for Evaluating Long-Term Dynamic ECG Analysis

Holtercare-Bench: A Multimodal Benchmark for Evaluating Long-Term Dynamic ECG Analysis

Holtercare-Bench:用于评估长期动态心电图分析的多模态基准

Abstract: While multimodal large language models (MLLMs) excel in medical applications, most of them favor static images or short-term signals. In the critical field of dynamic electrocardiograms (ECG), models struggle with complex temporal reasoning and diagnostic report generation due to a lack of high-quality datasets and benchmarks.

摘要: 尽管多模态大语言模型(MLLMs)在医疗应用中表现出色,但大多数模型更倾向于处理静态图像或短时信号。在动态心电图(ECG)这一关键领域,由于缺乏高质量的数据集和基准,模型在处理复杂的时间推理和诊断报告生成时往往表现不佳。

To address this, we introduce (i) Holtercare-23K, a large-scale multimodal dynamic ECG dataset comprising 22,980 QA pairs derived from 788 clinical Holter records and featuring a novel signal-video-text tri-modal alignment. Based on this dataset, we present (ii) Holtercare-Bench, a multimodal benchmark that evaluates models on temporal localization, clinical diagnosis, and global summarization.

为了解决这一问题,我们引入了 (i) Holtercare-23K,这是一个大规模多模态动态心电图数据集,包含从 788 份临床动态心电图记录中提取的 22,980 对问答(QA),并采用了新颖的“信号-视频-文本”三模态对齐方式。基于该数据集,我们提出了 (ii) Holtercare-Bench,这是一个多模态基准,用于评估模型在时间定位、临床诊断和全局总结方面的能力。

Zero-shot evaluations of leading MLLMs reveal a significant performance gap in processing ultra-long pathological sequences. However, fine-tuning representative models yields substantial improvements. This work illuminates the limitations of current MLLMs in electrophysiology and provides a foundational benchmark for long-term medical MLLMs. Our project is available at this https URL.

对领先的 MLLMs 进行的零样本(Zero-shot)评估显示,它们在处理超长病理序列时存在显著的性能差距。然而,对代表性模型进行微调后,性能得到了大幅提升。这项工作揭示了当前 MLLMs 在电生理学领域的局限性,并为长期医疗 MLLMs 提供了一个基础性基准。我们的项目地址请见:此链接