DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings

DrawingVQA:面向施工图纸多深度视觉-文本推理的真实世界基准

We introduce DrawingVQA, the first benchmark designed to evaluate multimodal large language models (MLLMs) on real-world construction drawings — a core media in architecture, civil, and many other engineering practices. 我们推出了 DrawingVQA,这是首个旨在评估多模态大语言模型(MLLMs)在真实施工图纸上表现的基准测试。施工图纸是建筑、土木及许多其他工程实践中的核心媒介。

Unlike natural images or schematic floor plans, construction drawings fuse abstract geometry, symbolic notation, tabular data, annotations, and domain-specific text, forming a uniquely complex visual-textual domain core to engineering workflows. 与自然图像或示意性平面图不同,施工图纸融合了抽象几何图形、符号标注、表格数据、注释以及特定领域的文本,形成了一个对工程工作流程至关重要的、极其复杂的视觉-文本领域。

DrawingVQA bridges this gap with 33 “Issued for Construction” drawings and 92 expertly curated question-answer pairs, spanning three reasoning depths: perceptual understanding, contextual interpretation, and domain-expert reasoning. DrawingVQA 通过 33 张“施工发布版”(Issued for Construction)图纸和 92 对由专家精心策划的问答对填补了这一空白,涵盖了三个推理深度:感知理解、上下文解读以及领域专家推理。

To evaluate model capabilities, we present a dual categorization framework to jointly analyze performance across seven construction-engineering and four MLLM capability dimensions — the first to explicitly map engineering workflows to AI reasoning competencies. 为了评估模型能力,我们提出了一个双重分类框架,用于联合分析模型在七个施工工程维度和四个 MLLM 能力维度上的表现——这是首个将工程工作流程明确映射到人工智能推理能力的框架。

Evaluations of state-of-the-art MLLMs reveal a substantial gap between model and expert performance, particularly at higher reasoning depths. 对当前最先进 MLLMs 的评估显示,模型表现与专家水平之间存在巨大差距,尤其是在更高深度的推理任务中。

This benchmark lays a foundation for domain-specialized multimodal reasoning to allow for advancement on integration of AI-driven understanding and real-world engineering workflows. 该基准为领域专业化的多模态推理奠定了基础,旨在推动人工智能驱动的理解能力与真实工程工作流程的深度融合。