XRF-to-Optical Field-of-View Localization with Vision Language Models
XRF-to-Optical Field-of-View Localization with Vision Language Models
基于视觉语言模型的 XRF 到光学视场定位
Abstract: Registering images acquired with different microscopy modalities is essential for relating complementary measurements of the same specimen. In correlative X-ray fluorescence (XRF) and optical microscopy, the XRF map often covers only a small region of an optical image acquired from the same or an adjacent tissue section. Field-of-view (FOV) localization is necessary but can be difficult when appearance and structure differ across modalities.
摘要: 对通过不同显微成像模态获取的图像进行配准,对于关联同一标本的互补测量结果至关重要。在相关性 X 射线荧光 (XRF) 和光学显微镜成像中,XRF 图谱通常仅覆盖从同一组织切片或相邻组织切片获取的光学图像的一小部分。视场 (FOV) 定位是必要的,但当不同模态之间的外观和结构存在差异时,定位往往非常困难。
Here we evaluate training-free vision language model (VLM) localization on two datasets representing same-section high-correspondence and adjacent-section low-correspondence imaging. We test unconstrained and metadata-constrained search and compare VLMs with geometric controls, classical template matching, and two alternative training-free approaches (DINOv2 and multiGradICON).
在此,我们评估了无需训练的视觉语言模型 (VLM) 在两个数据集上的定位能力,这两个数据集分别代表了同一切片的高对应度成像和相邻切片的低对应度成像。我们测试了无约束和元数据约束的搜索,并将 VLM 与几何控制、经典模板匹配以及两种无需训练的替代方法(DINOv2 和 multiGradICON)进行了比较。
Direct VLM prompting produced content-dependent spatial signals but was not reliable alone. Classical matching was most accurate when cross-modal structure was preserved but failed in the low-correspondence collection. A proposal-and-verify workflow used repeated VLM predictions as candidates and image-based similarity to select the final location. This workflow recovered useful localization in the low-correspondence regime.
直接使用 VLM 提示词可以产生依赖于内容的空间信号,但仅靠此方法并不可靠。当跨模态结构得以保留时,经典匹配方法最为准确,但在低对应度的数据集中则会失效。我们采用了一种“提议与验证”的工作流程,即利用重复的 VLM 预测作为候选结果,并结合基于图像的相似度来选择最终位置。该工作流程在低对应度的情况下实现了有效的定位。