CoVLM-Bench: A Real-World Benchmark for Cooperative Driving Question Answering and Planning
CoVLM-Bench: A Real-World Benchmark for Cooperative Driving Question Answering and Planning
CoVLM-Bench:用于协同驾驶问答与规划的真实世界基准
Vision-language models (VLMs) have made substantial progress in autonomous driving, but their success has primarily been studied in ego-centric scenes. Infrastructure-side observations provide views beyond the ego vehicle’s field of view, yet conventional cooperative-driving systems typically transform them into geometric representations for downstream perception and planning. Directly incorporating these views into VLMs offers an opportunity to improve cooperative scene understanding and trajectory planning. 视觉语言模型(VLMs)在自动驾驶领域取得了实质性进展,但其成功主要局限于以自我车辆为中心的场景。基础设施端的观测视角提供了超出自动驾驶车辆视野范围的信息,然而传统的协同驾驶系统通常将其转化为几何表示,用于后续的感知和规划。将这些视角直接整合到 VLM 中,为提升协同场景理解和轨迹规划能力提供了契机。
However, question answering and trajectory planning have not been jointly evaluated on the same real-world vehicle-infrastructure scenes. We present CoVLM-Bench, a benchmark for cooperative driving question answering (CDQA) and cooperative planning (CP) on vehicle-infrastructure paired scenes. CoVLM-Bench provides scene-grounded CDQA annotations, three-part rationales as auxiliary supervision, and future trajectory targets derived from recorded ego motion. 然而,目前尚未在相同的真实世界车路协同场景中对问答和轨迹规划进行联合评估。我们提出了 CoVLM-Bench,这是一个针对车路协同场景下的协同驾驶问答(CDQA)和协同规划(CP)的基准测试。CoVLM-Bench 提供了基于场景的 CDQA 标注、作为辅助监督的三部分推理逻辑,以及从记录的车辆运动中导出的未来轨迹目标。
It contains 2,196 paired frames with 35,136 CDQA annotations, while CP predicts six waypoints over a three-second horizon. The annotations combine model-assisted drafting, record-based computation, and human verification. Built upon CoVLM-Bench, we introduce CoVLM-Drive, a unified VLM baseline that directly uses paired views for both CDQA and CP. 该基准包含 2,196 帧配对图像和 35,136 条 CDQA 标注,其中 CP 任务预测三秒时间跨度内的六个航点。这些标注结合了模型辅助草拟、基于记录的计算以及人工验证。基于 CoVLM-Bench,我们引入了 CoVLM-Drive,这是一个统一的 VLM 基线模型,可直接利用配对视角同时进行 CDQA 和 CP 任务。
Experiments show that CDQA adaptation improves answer accuracy and that CoVLM-Drive reaches a lower FDE than the compared V2X planners; QA initialization and rationale supervision each reduce planning error. Together, CoVLM-Bench and CoVLM-Drive support the training and comparison of VLMs for cooperative scene understanding and planning. 实验表明,CDQA 的适配提升了回答准确率,且 CoVLM-Drive 在最终位移误差(FDE)指标上优于对比的 V2X 规划器;QA 初始化和推理逻辑监督均能有效降低规划误差。CoVLM-Bench 和 CoVLM-Drive 共同为协同场景理解与规划的 VLM 训练及对比提供了有力支持。