GaugeVLM: Structuring Spatial Supervision with Measured Geometric Interventions
GaugeVLM: Structuring Spatial Supervision with Measured Geometric Interventions
GaugeVLM:通过测量几何干预构建空间监督
Abstract: Vision-language models (VLMs) can contradict themselves across views of the same spatial relation and fail to respond when that relation changes. Addressing these failures requires supervision that captures error magnitude and geometric dependencies across observations, both of which remain implicit in training on individual answers or ordinal preferences.
摘要: 视觉语言模型(VLM)在面对同一空间关系的多个视角时,往往会产生自相矛盾的结论,且在空间关系发生变化时无法做出正确响应。要解决这些缺陷,需要一种能够捕捉观测结果中误差幅度和几何依赖关系的监督机制,而目前的训练方法仅依赖于单一答案或序数偏好,这些关键信息在训练中仍处于隐性状态。
Therefore, we introduce GaugeVLM, which makes this structure explicit through controlled object and camera interventions in explicit 3D scenes, producing linked observations with measured differences between spatial relations and shared truths across views. To translate this structure into learning signals, its core objective, GaugeDPO, converts measured errors into preference margins, directly supervises correct canonical rankings across views, and links intervention-induced answer-odds contrasts to measured relation changes with view-specific scales.
因此,我们引入了 GaugeVLM。该模型通过在显式 3D 场景中进行受控的对象和相机干预,使这种空间结构变得明确,从而生成具有关联性的观测数据,并测量不同视角下空间关系与共享真值之间的差异。为了将这种结构转化为学习信号,其核心目标函数 GaugeDPO 将测量到的误差转化为偏好边际,直接监督跨视角的正确规范排序,并将干预引起的答案概率对比与特定视角的空间关系变化联系起来。
Our analysis bounds canonical prediction error and establishes that the cross-view and intervention constraints can be jointly satisfied. Empirically, GaugeVLM improves all 10 established spatial metrics over supervised fine-tuning across three VLM backbones, with the main 7B model gaining 15.0 and 18.9 percentage points on MSMU distance and QSpatial+ respectively. These gains also extend to autonomous driving and embodied reasoning, demonstrating the robust generalization across domains.
我们的分析界定了规范预测误差,并证明了跨视角约束和干预约束可以同时得到满足。实验结果表明,在三个 VLM 主干网络上,GaugeVLM 在所有 10 项既定空间指标上均优于监督微调方法;其中,7B 参数的主模型在 MSMU 距离和 QSpatial+ 指标上分别提升了 15.0 和 18.9 个百分点。这些性能提升同样延伸到了自动驾驶和具身推理领域,证明了该模型在不同领域间具有强大的泛化能力。