World Editing: Intervening on Executable Worlds at Increasing Depth

World Editing: Intervening on Executable Worlds at Increasing Depth

世界编辑:在不断加深的深度上干预可执行世界

Abstract: Interactive world models are increasingly capable of generating environments and acting within them, yet deliberately editing an existing executable world remains underexplored. We formulate world editing as intervening on an existing world while preserving properties that should remain unchanged, and introduce intervention depth as an axis describing how strongly an edit couples world entities, dynamics, and systems.

摘要: 交互式世界模型在生成环境和在其中执行动作方面的能力日益增强,但对现有可执行世界进行刻意编辑的研究仍未得到充分探索。我们将“世界编辑”定义为在干预现有世界的同时,保持那些不应改变的属性,并引入了“干预深度”这一维度,用以描述编辑操作在多大程度上耦合了世界实体、动态和系统。

We instantiate this capability through industry-grade game modding and introduce IGMWorld, together with IGMBench, a benchmark of 110 tasks and over 1.1K executable state and behavioral criteria across Minecraft and Terraria. The tasks span property, entity, dynamics, and system interventions and are evaluated through deterministic executability, behavioral, preservation, and visual checks.

我们通过工业级游戏模组(modding)实现了这一能力,并推出了 IGMWorld,以及包含 110 个任务和超过 1.1K 个可执行状态与行为准则的基准测试集 IGMBench,涵盖了《我的世界》(Minecraft)和《泰拉瑞亚》(Terraria)。这些任务跨越了属性、实体、动态和系统干预,并通过确定性可执行性、行为、保持性及视觉检查进行评估。

Frontier coding agents already exhibit substantial world-editing capability: the strongest configuration solves 78.2% of tasks under a strict task-level criterion, while criterion-level performance reaches 94.8%. Reliability generally decreases with intervention depth, and this pattern persists even among tasks with similar numbers of evaluation criteria.

前沿编程智能体已经展现出相当可观的世界编辑能力:最强的配置在严格的任务级标准下解决了 78.2% 的任务,而在准则级表现上达到了 94.8%。可靠性通常随干预深度的增加而降低,即使在具有相似评估准则数量的任务中,这种模式依然存在。

Most failed edits still build and load successfully, suggesting that the main difficulty is making the edited world behave as requested. Visual consistency remains a separate weakness, with all evaluated configurations below 50% joint visual pass rate. These results show that world editing is a distinct capability from world generation and interaction, and that executable games provide a practical testbed for studying it.

大多数失败的编辑操作仍然能够成功构建和加载,这表明主要的困难在于使编辑后的世界按照预期运行。视觉一致性仍然是一个独立的弱点,所有评估配置的联合视觉通过率均低于 50%。这些结果表明,世界编辑是一项不同于世界生成和交互的独特能力,而可执行游戏为研究这一课题提供了实用的测试平台。