SciReC: Diagnostic Evaluation of Multimodal, Multi-Turn Relational Reasoning with Adaptive Interaction
SciReC: Diagnostic Evaluation of Multimodal, Multi-Turn Relational Reasoning with Adaptive Interaction
SciReC:基于自适应交互的多模态多轮关系推理诊断评估
Abstract: Relational reasoning requires the process of perceptual understanding, comparing, and integrating the underlying relationships between concepts. This ability consists of multiple categories, such as analogical, structural, and cause-effect, each capturing a different aspect of higher-order understanding.
摘要: 关系推理需要对概念间的潜在关系进行感知理解、比较和整合。这种能力包含多个类别,如类比、结构和因果关系,每一类都捕捉了高阶理解的不同侧面。
To examine the performance of multimodal large language models (MLLM) on these relational inference tasks, we developed SciReC, a model-adaptive multimodal academic dialog benchmark. As the relational reasoning process involves multiple representations and various factors (visual understanding, exhibiting knowledge, and memory recall), we propose DMRA, a deficit-based diagnostic framework that quantifies the contribution of these components to identify the primary cause of unsuccessful cases.
为了检验多模态大语言模型(MLLM)在这些关系推理任务上的表现,我们开发了 SciReC,这是一个模型自适应的多模态学术对话基准。由于关系推理过程涉及多种表征和多种因素(视觉理解、知识展示和记忆回溯),我们提出了 DMRA,这是一个基于缺陷的诊断框架,通过量化这些组件的贡献来识别推理失败案例的主要原因。
Claude 4.6 achieved the best performance on the overall relational score with 73%, followed by GPT 5.4 with 68%. Performance trends indicate that open-source models achieve their lowest scores on spatial relations, while proprietary models struggle more with hierarchical and sequential relations. Across domains, model performance is lowest on Astronomy and highest on Psychology.
Claude 4.6 在整体关系得分上表现最佳,达到 73%,其次是 GPT 5.4,得分为 68%。性能趋势表明,开源模型在空间关系上的得分最低,而闭源(专有)模型在层次和序列关系上表现更为吃力。在不同领域中,模型在天文学领域的表现最差,在心理学领域的表现最好。
The results of DMRA reveal that relational reasoning is the primary source of error across all models, followed by memory limitations.
DMRA 的结果显示,关系推理是所有模型产生错误的主要来源,其次是记忆限制。