Medical Image Alignment Assessment as a Test of Generalist Visual Reasoning in Frontier Multimodal Models

Frontier multimodal large language models (MLLMs) are increasingly positioned as general purpose visual reasoners as part of the quest for artificial general intelligence. A key test of this generality is whether they can perform novel visual judgments that humans can make reliably from visual evidence and task instructions, without task-specific parameter optimisation.

前沿多模态大语言模型(MLLMs)正日益被定位为通用视觉推理器,作为追求通用人工智能的一部分。检验这种通用性的一个关键标准是,它们是否能够在无需特定任务参数优化的情况下,仅凭视觉证据和任务指令,执行人类能够可靠做出的新颖视觉判断。

We investigate this question through the task of medical image alignment assessment, where the goal is to establish whether there is anatomical correspondence between two images. Human visual assessment of image alignment is still the gold standard and most common approach; however, it requires trained operators and is impractical to scale for large datasets.

我们通过医学图像对齐评估任务来研究这一问题,其目标是确定两幅图像之间是否存在解剖学上的对应关系。人类对图像对齐的视觉评估仍然是黄金标准和最常用的方法;然而,它需要受过培训的操作员,且难以扩展到大型数据集。

We evaluate recent generations of MLLMs on two exemplar medical image alignment tasks, varying both prompting strategies and image-presentation methods. We compare against a locally fine-tuned MLLM and a task-specific CNN to examine the trade-off between frontier general purpose models and smaller models that require specific task optimisation but can be used locally.

我们评估了近期几代 MLLMs 在两个医学图像对齐示例任务上的表现,并改变了提示策略和图像呈现方式。我们将结果与本地微调的 MLLM 和特定任务的 CNN 进行了比较,以考察前沿通用模型与需要特定任务优化但可在本地使用的较小模型之间的权衡。

We show that are reaching an inflection point, where frontier MLLMs can now perform effective visual assessment of medical image alignment. Models released only a few months ago generalise poorly and, in some settings, perform barely above chance, whereas GPT-6 achieves over 85% across almost all scenarios tested.

我们展示了目前正处于一个转折点,前沿 MLLMs 现在已经能够对医学图像对齐进行有效的视觉评估。仅在几个月前发布的模型泛化能力较差,在某些设置下表现仅略高于随机水平,而 GPT-6 在几乎所有测试场景中均达到了 85% 以上的准确率。

Fine-tuned local models can match or exceed frontier-model performance on the tasks on which they are trained, but transfer substantially less effectively to unseen settings. These findings identify medical image alignment as a useful test bed for generalist visual reasoning and suggest that frontier multimodal models are beginning to acquire capabilities that could support a common quality-control mechanism across heterogeneous medical-imaging pipelines.

经过微调的本地模型在训练任务上可以达到或超过前沿模型的性能,但在未见过的场景中迁移效果明显较差。这些发现表明,医学图像对齐是通用视觉推理的一个有用的测试平台,并暗示前沿多模态模型正开始获得能够支持异构医学成像流程中通用质量控制机制的能力。