Octrees as an Explicit 3D Language
Octrees as an Explicit 3D Language
Abstract: Existing 3D large language models (LLMs) compromise on two fronts: they compress shapes into latent codebook indices or coordinate text, which removes spatial structure from what the model observes, and they acquire the 3D modality by fine-tuning the backbone, which overwrites its general language ability.
摘要: 现有的 3D 大语言模型(LLM)在两个方面存在妥协:它们将形状压缩为潜在码本索引或坐标文本,这抹去了模型所观察到的空间结构;同时,它们通过微调主干网络来获取 3D 模态,这会覆盖其原有的通用语言能力。
We present OctLLM, which addresses both limitations. Geometry enters as an explicit 3D sequence of octree occupancy tokens. However, full octree sequences grow rapidly with depth; OctLLM therefore randomly empties penultimate-level nodes and omits descendants while preserving shape, yielding a shorter coordinate- and depth-anchored Sparse Octree (S-Octree) for position-aware mask-modeling generation and 3D understanding.
我们提出了 OctLLM,旨在解决上述两个局限性。几何信息以八叉树占用标记(octree occupancy tokens)的显式 3D 序列形式输入。然而,完整的八叉树序列会随深度迅速增长;因此,OctLLM 在保留形状的同时,随机清空倒数第二层的节点并省略其后代,从而生成了一种更短的、具有坐标和深度锚定的稀疏八叉树(S-Octree),用于位置感知的掩码建模生成和 3D 理解。
On the other front, existing methods introduce a new modality with full fine-tuning or LoRA, but full fine-tuning is costly, LoRA limits 3D capacity, and both modify the language pathway. OctLLM instead adds 3D capacity in parameters separate from the pretrained ones: mesh tokens are routed through independent trainable branches in a subset of blocks while text and image tokens retain the frozen vision-language pathway, and the two streams interact through shared self-attention.
在另一方面,现有方法通过全量微调或 LoRA 引入新模态,但全量微调成本高昂,LoRA 则限制了 3D 容量,且两者都会修改语言路径。相比之下,OctLLM 在预训练参数之外增加了独立的 3D 容量参数:网格标记(mesh tokens)通过部分区块中的独立可训练分支进行路由,而文本和图像标记则保留冻结的视觉-语言路径,两者通过共享的自注意力机制进行交互。
It trains far fewer parameters than full fine-tuning, yet sets a new state of the art among unified multimodal LLMs, lowering image-to-3D FID by $17.4%$ and raising render-grounded captioning by $28.7$ points over ShapeLLM-Omni, while matching the backbone on general language benchmarks.
与全量微调相比,该模型训练的参数量大幅减少,却在统一多模态 LLM 中树立了新的技术标杆。相比 ShapeLLM-Omni,其图像到 3D 的 FID 指标降低了 $17.4%$,基于渲染的图像描述(render-grounded captioning)得分提升了 $28.7$ 分,同时在通用语言基准测试中保持了与主干模型相当的性能。