SCION: Scene Composition with Instanced Neural Primitives
SCION: Scene Composition with Instanced Neural Primitives
SCION:基于实例化神经图元的场景合成
Real-world scenes are compositional: bricks, blades of grass, pebbles, and tree leaves recur across human-built and natural environments. Existing neural scene representations model these elements independently. Most 3D Gaussian Splatting and follow-up abstraction and compression methods treat each element as unique, fitting millions of independent Gaussians per scene. Prior methods like Splat and Replace fit template objects, but they require mostly manual selection of repeated elements. As a result, these representations store redundant parameters and provide weak manipulation handles for downstream tasks.
现实世界的场景具有组合性:砖块、草叶、鹅卵石和树叶在人造和自然环境中反复出现。现有的神经场景表示法将这些元素独立建模。大多数 3D 高斯泼溅(3D Gaussian Splatting)及其后续的抽象和压缩方法都将每个元素视为唯一的,为每个场景拟合数百万个独立的高斯分布。诸如“Splat and Replace”之类的先前方法虽然拟合了模板对象,但它们大多需要手动选择重复元素。因此,这些表示法存储了冗余参数,且为下游任务提供的操作控制能力较弱。
We introduce SCION, a hierarchical compositional scene representation that replaces independent Gaussians with a compact vocabulary of reusable primitives and lightweight world-space instances that place transformed copies throughout the scene. We fit this representation to multi-view captures via a joint optimization over discrete and continuous scene parameters, combining two-level densification over splats and instances with an adversarial loss that preserves detail across shared primitives.
我们引入了 SCION,这是一种分层组合式场景表示法。它用一套紧凑的可重用图元词汇表,以及在场景中放置变换副本的轻量级世界空间实例,取代了独立的高斯分布。我们通过对离散和连续场景参数进行联合优化,将这种表示法拟合到多视图捕捉数据中,并结合了针对泼溅和实例的两级致密化处理,以及一种在共享图元间保持细节的对抗性损失函数。
The recovered structure yields a compact, controllable representation while maintaining high quality even at 1.2 MB. SCION achieves rate-distortion favorable to existing Gaussian compression methods, and it enables instance-level scene editing and animation without retraining. Our results show that neural scene representations need not memorize scenes as independent primitives; they can discover reusable parts.
恢复后的结构不仅产生了一种紧凑且可控的表示,即使在 1.2 MB 的大小下也能保持高质量。SCION 在速率-失真(rate-distortion)表现上优于现有的高斯压缩方法,并且无需重新训练即可实现实例级的场景编辑和动画制作。我们的结果表明,神经场景表示法无需将场景记忆为独立的图元,它们完全可以发现可重用的部分。