AerialDojo-200K: A Large-Scale Benchmark Suite for Open-World Aerial Object-Goal Search

AerialDojo-200K: A Large-Scale Benchmark Suite for Open-World Aerial Object-Goal Search

AerialDojo-200K:用于开放世界空中目标搜索的大规模基准套件

Abstract: Open-world aerial object-goal search is a foundational yet challenging task, requiring aerial agents to autonomously explore large-scale, unstructured three-dimensional environments and reach target objects specified by semantic descriptions or reference images, rather than following route-specific instructions. However, research in this task remains at a nascent stage and relies on small, environment-specific benchmarks with heterogeneous action spaces and data formats. These limitations hinder large-scale training and cross-benchmark evaluation, constraining the scalability and generalizability of aerial agents.

摘要: 开放世界空中目标搜索是一项基础且具有挑战性的任务,它要求空中智能体能够自主探索大规模、非结构化的三维环境,并根据语义描述或参考图像找到目标物体,而不是仅仅遵循特定的路线指令。然而,该领域的研究仍处于起步阶段,且依赖于规模较小、环境特定且动作空间与数据格式各异的基准测试。这些局限性阻碍了大规模训练和跨基准评估,限制了空中智能体的可扩展性和泛化能力。

To address this problem, we propose AerialDojo-200K, a large-scale benchmark suite for open-world aerial object-goal search, with 3 times as many scenes and 18.7 times as many task instances as the largest existing benchmark for this task. Specifically, we construct 42 simulation scenes spanning four scene families and 21 scene types, including 18 urban, 12 natural, six infrastructure, and six disaster scenes. To ensure data quality, 12 annotators spent two months manually annotating 109 landmarks, 2099 target objects, and 2099 object anchors across these scenes.

为了解决这一问题,我们提出了 AerialDojo-200K,这是一个用于开放世界空中目标搜索的大规模基准套件,其场景数量是现有最大同类基准的 3 倍,任务实例数量则是其 18.7 倍。具体而言,我们构建了 42 个模拟场景,涵盖了 4 个场景系列和 21 种场景类型,包括 18 个城市场景、12 个自然场景、6 个基础设施场景和 6 个灾难场景。为确保数据质量,12 名标注员耗时两个月,手动标注了这些场景中的 109 个地标、2099 个目标物体以及 2099 个物体锚点。

We further construct 205,732 task instances, comprising over 100K semantic-goal and over 100K image-goal instances across Base, Standard, and Long-Horizon settings. Each task instance includes a collision-free reference trajectory and corresponding multi-view video recordings. We also develop a unified evaluation framework with a scene partition comprising 21 in-distribution scenes and 21 out-of-distribution scenes. Finally, our evaluation of five open-source and four closed-source multimodal large language models reveals that there is still a long way to go toward achieving general-purpose aerial agents.

我们进一步构建了 205,732 个任务实例,包含超过 10 万个语义目标实例和超过 10 万个图像目标实例,涵盖了基础(Base)、标准(Standard)和长视距(Long-Horizon)三种设置。每个任务实例都包含一条无碰撞的参考轨迹及相应的多视角视频记录。我们还开发了一个统一的评估框架,将场景划分为 21 个分布内(in-distribution)场景和 21 个分布外(out-of-distribution)场景。最后,我们对 5 个开源和 4 个闭源多模态大语言模型进行了评估,结果表明,要实现通用型空中智能体,仍有很长的路要走。