Axolotl3D: a Unified Framework for Faithful 3D Shape Completion
Axolotl3D: a Unified Framework for Faithful 3D Shape Completion
Axolotl3D:用于高保真 3D 形状补全的统一框架
Abstract: Recent 3D generative models produce high-quality geometry from a single image using large-scale priors and diffusion architectures. However, they assume complete visibility and single-view inputs, limiting applicability in multi-view, occluded, or editing scenarios. 摘要: 近期的 3D 生成模型利用大规模先验知识和扩散架构,能够从单张图像中生成高质量的几何形状。然而,这些模型通常假设物体完全可见且仅有单视图输入,这限制了其在多视图、遮挡或编辑场景中的应用。
Although prior works address these challenges individually, they lack a unified framework for controllable 3D completion under diverse conditioning signals. We present Axolotl3D, a multi-modal and occlusion-aware 3D generation model that jointly conditions on images, visibility masks, camera parameters, and a partial point cloud. 尽管先前的研究已分别解决了这些挑战,但仍缺乏一个能够在多种条件信号下实现可控 3D 补全的统一框架。我们提出了 Axolotl3D,这是一个多模态且具备遮挡感知能力的 3D 生成模型,它能够联合利用图像、可见性掩码、相机参数以及部分点云作为条件输入。
The point cloud serves as a geometric anchor promoting faithful shape completion, while camera parameters ensure consistent multi-view alignment in a shared 3D coordinate system. A unified training strategy synthesizes diverse conditioning regimes from large-scale 3D data, enabling robust cross-modal reasoning. 点云作为几何锚点,促进了高保真的形状补全,而相机参数则确保了在共享 3D 坐标系中多视图对齐的一致性。统一的训练策略从大规模 3D 数据中综合了多种条件机制,从而实现了稳健的跨模态推理。
Experiments on Toys4K and OmniObject3D demonstrate state-of-the-art performance under both clean and occluded settings, as well as strong results in real-world reconstruction and geometry-consistent editing. 在 Toys4K 和 OmniObject3D 数据集上的实验表明,该模型在清晰和遮挡环境下均达到了最先进的性能,并在真实世界重建和几何一致性编辑方面表现出色。