Systematic Multi-Agent Vision-and-Language Navigation: Formulation, Benchmark, and Method

Systematic Multi-Agent Vision-and-Language Navigation: Formulation, Benchmark, and Method

系统化多智能体视觉语言导航:形式化、基准测试与方法

Vision-and-Language Navigation (VLN) has largely focused on a single agent following a single instruction, yet many real-world applications require teams of robots to tackle tasks beyond the capabilities of any individual agent. 视觉语言导航(VLN)的研究主要集中在单个智能体遵循单条指令上,然而许多现实世界的应用需要机器人团队协作,以完成超出单个智能体能力范围的任务。

We present Systematic Multi-Agent Vision-and-Language Navigation, providing, to our knowledge, the first systematic formalization of multi-agent VLN as a constrained coordination problem: each mission consists of subtasks carrying dependency and resource constraints (presence locks and holding chains). 我们提出了“系统化多智能体视觉语言导航”(Systematic Multi-Agent Vision-and-Language Navigation),据我们所知,这是首次将多智能体 VLN 系统地形式化为一种受约束的协调问题:每项任务由带有依赖关系和资源约束(存在锁和持有链)的子任务组成。

A verified four-stage crafting pipeline instantiates the task as MAVLN, comprising 11,724 episodes across 145 scenes with teams of up to four agents under three instruction regimes, accompanied by tailored constraint-aware metrics. 我们通过一个经过验证的四阶段构建流程将该任务实例化为 MAVLN,包含 145 个场景中的 11,724 个片段,支持最多四个智能体的团队在三种指令模式下进行协作,并配备了专门的约束感知评估指标。

We further present TRISS, a coordination-ready navigation system coupling an LLM-based subtask scheduler, a shared topological memory that turns each agent’s exploration into team knowledge, and a conflict-aware execution mechanism that realizes simultaneous intentions as collision-free routes. 此外,我们还提出了 TRISS,这是一个具备协调能力的导航系统。它结合了基于大语言模型(LLM)的子任务调度器、将每个智能体的探索转化为团队知识的共享拓扑记忆,以及能够将并发意图转化为无碰撞路径的冲突感知执行机制。

Extensive experiments establish TRISS as a comprehensive baseline and reveal substantial room for improvement across scheduling, planning, and execution, highlighting the challenges of coordinating under MAVLN task constraints. 广泛的实验表明,TRISS 是一个全面的基准模型,同时也揭示了在调度、规划和执行方面仍有巨大的改进空间,凸显了在 MAVLN 任务约束下进行协调所面临的挑战。