SPW-Nav: A Streaming Panoramic World Model for Language-Guided Navigation

Computer Science > Computer Vision and Pattern Recognition arXiv:2610.08941 (cs) [Submitted on 6 Oct 2026] Title: SPW-Nav: A Streaming Panoramic World Model for Language-Guided Navigation Authors: Yunheng Liu, Ziqi Cai, Siqi Yang, Yimu Wang, Minggui Teng, Jiaming Tan, Shuchen Weng, Erwin Wu, Kaipeng Zhang, Boxin Shi.

计算机科学 > 计算机视觉与模式识别 arXiv:2610.08941 (cs) [提交于 2026 年 10 月 6 日] 标题:SPW-Nav:一种用于语言引导导航的流式全景世界模型 作者:Yunheng Liu, Ziqi Cai, Siqi Yang, Yimu Wang, Minggui Teng, Jiaming Tan, Shuchen Weng, Erwin Wu, Kaipeng Zhang, Boxin Shi。

Abstract: Language-guided panoramic video generation benefits various downstream applications, such as interactive 3D scene exploration, virtual reality experiences, and embodied agent training. Existing panoramic generators follow predefined trajectories, and interactive world models act through low-level actions in perspective views.

摘要:语言引导的全景视频生成有益于各种下游应用,例如交互式 3D 场景探索、虚拟现实体验以及具身智能体训练。现有的全景生成器遵循预定义的轨迹,而交互式世界模型则通过透视视图中的低级动作进行操作。

We propose SPW-Nav, a streaming panoramic world model that understands movement instructions and streams one minute of 2K 360-degree video in real time from a single panorama. SPW-Nav interprets each instruction in the previously generated panorama as camera motion.

我们提出了 SPW-Nav,这是一种流式全景世界模型,它能够理解移动指令,并能从单张全景图实时流式传输一分钟的 2K 360 度视频。SPW-Nav 将先前生成的全景图中的每条指令解释为摄像机运动。

Spherical rotation decoupling applies rotation exactly on the sphere, pose-aligned conditioning keeps translation inputs bounded over long streams, and a multi-term memory with a few-step generator continues the scene as instructions change. We also build SPW-NavSet, panoramic videos with camera trajectories and verified instructions.

球形旋转解耦在球面上精确应用旋转,姿态对齐调节使平移输入在长流中保持受限,而具有少步生成器的多项记忆机制则随着指令的变化延续场景。我们还构建了 SPW-NavSet,即包含摄像机轨迹和已验证指令的全景视频集。

Driven by language, SPW-Nav outperforms prior panoramic generators in camera-following accuracy and video quality, and supports on-the-fly instruction switching.

在语言驱动下,SPW-Nav 在摄像机跟随准确性和视频质量方面优于先前的全景生成器,并支持即时指令切换。