Autonomous Driving Research Requires a Community-Driven Data Paradigm
Autonomous driving has made remarkable progress, with recent AI advances enabling commercial deployments that are reshaping urban mobility. Yet the field remains far from its universal social promise: autonomous systems that can operate robustly anywhere, anytime, for anyone.
自动驾驶已经取得了显著进展,近期人工智能的进步使得商业部署成为可能,并正在重塑城市交通。然而,该领域距离其普遍的社会承诺仍有很大差距:即实现能够在任何地点、任何时间、为任何人稳健运行的自动驾驶系统。
We posit that this gap is not merely a modeling problem, but a problem of the prevailing data paradigm. Current research relies heavily on a few benchmark datasets with limited spatial and scenario coverage, even though the community has collectively produced over 600 autonomous driving datasets across nearly 50 countries.
我们认为,这一差距不仅仅是一个建模问题,更是当前数据范式存在的问题。尽管学术界已在近 50 个国家共同产生了超过 600 个自动驾驶数据集,但目前的研究仍过度依赖少数几个在空间和场景覆盖上有限的基准数据集。
However, this abundance has not translated into broad research impact: most datasets remain significantly underused due to fragmentation, limited visibility, incompatible protocols, and benchmark incentives that concentrate attention on a few dominant datasets.
然而,这种丰富的数据资源并未转化为广泛的研究影响力:由于数据碎片化、可见度有限、协议不兼容,以及将注意力集中在少数主流数据集上的基准激励机制,大多数数据集仍未得到充分利用。
We therefore argue that autonomous driving research requires a collaborative, community-driven data paradigm. Such a paradigm would improve the discovery, reuse, integration, and evaluation of diverse datasets; make underexplored data easier and more rewarding to study; and lower the barrier for new contributors.
因此,我们主张自动驾驶研究需要一种协作式的、社区驱动的数据范式。这种范式将改善多样化数据集的发现、重用、整合与评估;使未被充分探索的数据更易于研究且更具价值;并降低新贡献者的参与门槛。
We outline its key principles, illustrate an early realization, and call for collaboration across academia and industry to transform fragmented datasets into shared community infrastructure for anytime-anywhere autonomy.
我们概述了其核心原则,展示了早期的实现方案,并呼吁学术界和工业界开展合作,将碎片化的数据集转化为共享的社区基础设施,以实现随时随地的自动驾驶。