DistScene: Object-to-Scene Distillation for 3D Scene Generation
Computer Science > Computer Vision and Pattern Recognition arXiv:2610.06960 (cs) [Submitted on 3 Oct 2026] Title:DistScene: Object-to-Scene Distillation for 3D Scene Generation Authors:Kunming Luo, Hongyu Yan, Ken Deng, Chengcheng Zhou, Tianyu Liu, Haipeng Li, Haibin Huang, Xuelong Li, Ping Tan.
计算机科学 > 计算机视觉与模式识别 arXiv:2610.06960 (cs) [提交于 2026 年 10 月 3 日] 标题:DistScene:用于 3D 场景生成的对象到场景蒸馏。作者:Kunming Luo, Hongyu Yan, Ken Deng, Chengcheng Zhou, Tianyu Liu, Haipeng Li, Haibin Huang, Xuelong Li, Ping Tan。
Abstract: We present DistScene, a framework for single-image compositional 3D scene generation by jointly modeling the environment and individual objects. Unlike existing methods that represent scenes primarily as collections of objects, we model the environment as an explicit scene component to provide geometric context for object placement.
摘要:我们提出了 DistScene,这是一个通过联合建模环境和单个对象来实现单图像组合式 3D 场景生成的框架。与现有的主要将场景表示为对象集合的方法不同,我们将环境建模为一个显式的场景组件,从而为对象放置提供几何上下文。
Specifically, we introduce Scene-Frame Generation, which jointly generates separate environment and object components in a shared coordinate frame, allowing their geometry and relative placement to be learned together. Then we introduce Object-Centric Refinement to refine each object in a local frame with scene context.
具体而言,我们引入了场景框架生成(Scene-Frame Generation),它在共享坐标系中联合生成独立的环境和对象组件,使得它们的几何形状和相对位置能够被共同学习。随后,我们引入了以对象为中心的细化(Object-Centric Refinement),在场景上下文的辅助下,在局部框架内对每个对象进行细化。
Finally, we develop Object-to-Scene Distillation to transfer pretrained object-generation priors to scene generation through automatically composed and rendered synthetic scenes. Evaluations on indoor and outdoor benchmarks demonstrate improved scene-level spatial coherence over the evaluated baselines.
最后,我们开发了对象到场景蒸馏(Object-to-Scene Distillation),通过自动合成和渲染的场景,将预训练的对象生成先验知识迁移到场景生成中。在室内和室外基准测试上的评估表明,与所评估的基准方法相比,该方法提高了场景级的空间连贯性。