One Geometry, Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models
One Geometry, Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models
同一几何,不同结果:视觉-语言模型中模态鸿沟的读出依赖效应
Abstract: Contrastive vision-language models learn shared embedding spaces by aligning matched image-text pairs, yet their representations remain separated by a modality gap. Prior work reports divergent effects of modifying this gap: reducing it can improve zero-shot classification and cross-modal alignment, whereas removing gap-related structure can degrade image-text retrieval.
摘要: 对比式视觉-语言模型通过对齐匹配的图像-文本对来学习共享嵌入空间,但它们的表示仍然被“模态鸿沟”(modality gap)所分隔。先前的研究报告了修改该鸿沟所带来的不同影响:减小鸿沟可以改善零样本分类和跨模态对齐,而去除与鸿沟相关的结构则可能降低图像-文本检索的性能。
In this paper, we provide a unified geometric explanation for these task-dependent effects. Across CLIP and SigLIP encoders, we find that a single dominant direction captures 94.4-99.9% of the squared norm of the image-text mean separation, revealing that the mean-separation component is approximately rank-one.
在本文中,我们为这些任务依赖效应提供了一种统一的几何解释。通过对 CLIP 和 SigLIP 编码器的研究,我们发现一个单一的主导方向捕获了图像-文本均值分离平方范数的 94.4%-99.9%,这揭示了均值分离分量近似为秩一(rank-one)。
A decomposition of the similarity score then identifies three task-specific roles. In zero-shot classification, query-side fixed gap-offset subtraction is exactly equivalent to an additive class bias. In standard cross-modal retrieval, projecting out the gap direction and renormalising residuals discards candidate-specific norm information, inducing a multiplicative ranking distortion; a geometry-derived exponent tracks the grid-search optimum (Spearman rho = 0.93) and restores performance in some settings, although the gains transfer unevenly.
随后,通过对相似度得分进行分解,我们确定了三种特定于任务的作用。在零样本分类中,查询侧固定的鸿沟偏移减法完全等同于一种加性类别偏差。在标准跨模态检索中,将鸿沟方向投影出去并对残差进行归一化会丢弃候选对象特定的范数信息,从而导致乘性排序失真;一个源自几何的指数可以追踪网格搜索的最优值(Spearman 相关系数 = 0.93),并在某些设置下恢复性能,尽管这些增益的迁移并不均匀。
In mixed-modal retrieval, the gap direction sorts candidates by modality; its removal can improve cross-modal ranking, unlike random or non-gap controls. Residual semantic structure after removal defines the limits of the rank-one account. Together, these results explain why gap modification can improve, degrade, or restore performance across downstream settings. By clarifying when and why gap modification changes model behavior, this account provides a principled basis for selecting gap interventions in similarity-based vision-language systems across evaluated downstream tasks.
在混合模态检索中,鸿沟方向按模态对候选对象进行排序;与随机或非鸿沟控制组不同,去除该方向可以改善跨模态排序。去除后的残余语义结构定义了秩一解释的局限性。总之,这些结果解释了为什么修改鸿沟可以在不同的下游设置中改善、降低或恢复性能。通过阐明修改鸿沟在何时以及为何会改变模型行为,本研究为在基于相似度的视觉-语言系统中选择鸿沟干预措施提供了原则性基础,并涵盖了所评估的各项下游任务。