EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling

EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling

EditHero:用于长程、部件级 3D 编辑与氛围建模的基准测试

Abstract: 3D editing methods are usually tested on a single edit, yet an asset is built through a long sequence of revisions, each of which must implement the requested change while leaving everything else unchanged.

摘要: 3D 编辑方法通常是在单次编辑任务上进行测试的,然而,一个资产的构建往往需要经过一系列长期的修订过程,其中每一次修订都必须在实现所需更改的同时,保持其余部分不变。

We introduce EditHero, to our knowledge the first benchmark for long-horizon, part-level 3D editing, with natural-language instructions and target images for both geometry and texture. A deterministic assembly engine produces the exact target after every edit, and every sequence is reviewed by hand.

我们推出了 EditHero,据我们所知,这是首个针对长程、部件级 3D 编辑的基准测试,它提供了涵盖几何形状和纹理的自然语言指令及目标图像。一个确定性的组装引擎会在每次编辑后生成精确的目标结果,且每一组编辑序列都经过了人工审核。

We use EditHero to compare 2 opposite approaches to 3D editing. Non-agentic methods operate top down, regenerating the object from a learned 3D representation and inferring what to keep. In contrast, LLM/VLM agents operate bottom up, editing through code that inspects the mesh and rewrites only the parts required by instructions.

我们利用 EditHero 对两种截然不同的 3D 编辑方法进行了比较。非智能体(Non-agentic)方法采用自上而下的方式,通过学习到的 3D 表示重新生成对象,并推断哪些部分需要保留。相比之下,LLM/VLM 智能体则采用自下而上的方式,通过代码进行编辑,这些代码会检查网格并仅重写指令所要求的部分。

The non-agentic methods often miss the requested change and disturb regions that should stay fixed. Most LLMs follow instructions more closely, and all of them preserve the unedited parts better, but each of their edits takes minutes. We will release the engine and the edit sequences to support research on reliable iterative 3D editing.

非智能体方法往往会遗漏所要求的更改,并干扰那些本应保持不变的区域。大多数大语言模型(LLM)能更紧密地遵循指令,且都能更好地保留未编辑的部分,但它们的每次编辑都需要耗费数分钟时间。我们将发布该引擎及编辑序列,以支持关于可靠迭代式 3D 编辑的研究。