PanoPed: Beyond Bounding Boxes for Sim-to-Real Panoramic Pedestrian Tracking
Computer Science > Computer Vision and Pattern Recognition arXiv:2610.08826 (cs) [Submitted on 25 Sep 2026] Title: PanoPed: Beyond Bounding Boxes for Sim-to-Real Panoramic Pedestrian Tracking Authors: Qinfeng Zhu, Weiguang Zhao, Yunxi Jiang, Anh Nguyen, Lei Fan.
计算机科学 > 计算机视觉与模式识别 arXiv:2610.08826 (cs) [提交于 2026 年 9 月 25 日] 标题:PanoPed:超越边界框的仿真到现实全景行人跟踪 作者:Qinfeng Zhu, Weiguang Zhao, Yunxi Jiang, Anh Nguyen, Lei Fan。
Abstract: Full-sphere panoramic cameras let fixed monitoring systems and mobile robots track people in every direction, but a planar bounding box does not fully describe where a person is on the sphere. We introduce PanoPed, a sim-to-real benchmark for pedestrian tracking on the full sphere. PanoPed-S contains 108,000 frames from fixed, quadruped-mounted, and drone-mounted cameras, with synchronized masks, depth, camera poses, and 3D pedestrian states.
摘要:全球面全景摄像机使固定监控系统和移动机器人能够跟踪各个方向的人员,但平面边界框无法完全描述一个人在球体上的位置。我们引入了 PanoPed,这是一个用于全球面行人跟踪的仿真到现实基准测试。PanoPed-S 包含来自固定式、四足机器人搭载式和无人机搭载式摄像机的 108,000 帧图像,并配有同步的掩码、深度、摄像机姿态和 3D 行人状态。
PanoPed-R adds 28,002 real frames from fixed cameras, 16,247 of them densely annotated. We find that an ERP rectangle cannot uniquely determine the spherical center and angular extent of the visible person, while the detector’s visual query still carries information about them. Inspired by the sextant’s use of angular measurements to locate objects, we propose Sextant, a plug-and-play angular localization head with only about 0.035M parameters.
PanoPed-R 增加了 28,002 帧来自固定摄像机的真实图像,其中 16,247 帧经过了密集标注。我们发现,等距柱状投影(ERP)矩形无法唯一确定可见人员的球心和角范围,而检测器的视觉查询仍包含有关它们的信息。受六分仪利用角度测量定位物体的启发,我们提出了 Sextant,这是一个即插即用的角度定位头,参数量仅约 0.035M。
It reuses a frozen detector, keeps track identities unchanged, and needs no extra image encoder. Sextant gives the best result in our PanoPed-S test comparison, raising the strongest baseline, MOTIP, from 47.30 to 49.49 HOTA, with gains on all eight test sequences. Without fine-tuning on real data, the same synthetic-trained heads improve MOTIP and HAT by 0.96-1.14 HOTA on real video, and both seeds improve every real sequence. HAT+Sextant scores best among the compared systems that add no localization image encoder.
它复用了冻结的检测器,保持跟踪身份不变,且无需额外的图像编码器。Sextant 在我们的 PanoPed-S 测试对比中给出了最佳结果,将最强基准 MOTIP 的 HOTA 从 47.30 提升至 49.49,并在所有八个测试序列上均有提升。在未对真实数据进行微调的情况下,相同的合成训练头在真实视频上将 MOTIP 和 HAT 的 HOTA 提升了 0.96-1.14,且两个种子模型都改善了每一个真实序列。在未添加定位图像编码器的对比系统中,HAT+Sextant 的得分最高。