TestHallVQA: Exploring LVLMs' Document-Level Reasoning under Redundant Contexts from Scientific Exams

TestHallVQA: Exploring LVLMs’ Document-Level Reasoning under Redundant Contexts from Scientific Exams

TestHallVQA:探索大型视觉语言模型在科学考试冗余上下文下的文档级推理能力

Large Vision—Language Models (LVLMs) are increasingly expected to perform visual question answering (VQA) over planar media. However, existing planar VQA benchmarks typically emphasize isolated challenges: some emphasize long-document understanding with limited reasoning depth, while others require complex visual reasoning but remain restricted to single-page, noise-free settings.

人们对大型视觉语言模型(LVLMs)在平面媒体上执行视觉问答(VQA)的期望日益增长。然而,现有的平面 VQA 基准测试通常侧重于孤立的挑战:一些侧重于长文档理解但推理深度有限,而另一些则需要复杂的视觉推理,但仍局限于单页、无噪声的场景。

Moreover, through theoretical analysis, we identify the impact of irrelevant visual tokens, which leads to measurable performance degradation but has received little attention with respect to systematic quantification.

此外,通过理论分析,我们确定了无关视觉标记(visual tokens)的影响,这会导致可衡量的性能下降,但在系统性量化方面却鲜受关注。

To address these limitations, we introduce TestHallVQA, a multi-image VQA benchmark that simultaneously embodies document-level scale and the difficulty of human examinations, while providing comprehensive task coverage.

为了解决这些局限性,我们引入了 TestHallVQA,这是一个多图像 VQA 基准测试,它同时体现了文档级规模和人类考试的难度,并提供了全面的任务覆盖。

Leveraging TestHallVQA’s ability to controllably inject multi-level contextual redundancy, we further propose a novel metric, F1-R\textsuperscript{2}, which jointly quantifies LVLMs’ computational reasoning capability and their evidence retrieval robustness against document-level redundancy.

利用 TestHallVQA 可控地注入多级上下文冗余的能力,我们进一步提出了一种新的指标 F1-R\textsuperscript{2},该指标共同量化了 LVLMs 的计算推理能力及其在面对文档级冗余时的证据检索鲁棒性。

Extensive experiments and analyses on mainstream LVLMs reveal their latent deficiencies across multiple dimensions, offering concrete insights and directions for future research. The associated datasets, code, and complete theoretical derivations are available at this https URL.

对主流 LVLMs 进行的大量实验和分析揭示了它们在多个维度上的潜在缺陷,为未来的研究提供了具体的见解和方向。相关的数据集、代码和完整的理论推导可在该链接获取。