DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents

DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents

DeskForge:面向计算机使用智能体的桌面环境密集监督

Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually similar controls compete for attention. Existing training data rarely pair such scenes with dense annotations or vary them in a controlled way. 计算机使用智能体需要在复杂的桌面场景中可靠地定位操作目标,而在这些场景中,多个应用程序、重叠的窗口以及视觉上相似的控件往往会相互干扰。现有的训练数据很少将此类场景与密集标注配对,也难以以受控的方式对其进行变化。

We introduce DeskForge, a controllable desktop environment that composes and explores real applications to generate large-scale supervision for computer-use agents. It varies application states, content, window layout, appearance, and resolution, and fuses screenshots, accessibility trees, and window geometry into dense element annotations while recording the outcome of each executed action. 我们推出了 DeskForge,这是一个可控的桌面环境,它通过组合和探索真实应用程序,为计算机使用智能体生成大规模的监督数据。它能够改变应用程序的状态、内容、窗口布局、外观和分辨率,并将屏幕截图、辅助功能树和窗口几何信息融合为密集的元素标注,同时记录每次执行操作的结果。

Using this environment, we construct DeskForge-1M, a corpus of 1.2M annotated desktop observations containing 159.7M element instances. We fine-tune four vision-language models on 200K grounding examples drawn from DeskForge-1M. All four improve across held-out desktop conditions and on all five external GUI grounding benchmarks; for Qwen3.5-4B, accuracy increases by 11.51 percentage points on ScreenSpot-Pro and 10.11 points on OSWorld-G. 利用该环境,我们构建了 DeskForge-1M,这是一个包含 120 万个带标注桌面观测数据的语料库,其中包含 1.597 亿个元素实例。我们在从 DeskForge-1M 中提取的 20 万个定位示例上对四个视觉语言模型进行了微调。所有四个模型在未见过的桌面条件下以及全部五个外部 GUI 定位基准测试中均有所提升;对于 Qwen3.5-4B 模型,其在 ScreenSpot-Pro 上的准确率提高了 11.51 个百分点,在 OSWorld-G 上提高了 10.11 个百分点。

The gains also translate to long-horizon task completion: under a fixed planner, the fine-tuned action models solve more WebArena-Infinity and OpenApps tasks, with Qwen3.5-4B increasing from 31 to 50 of 119 tasks and from 3 to 15 of 100 tasks, respectively. These results show that controllable composition of real desktop environments provides a scalable source of supervision for improving both GUI grounding and long-horizon computer use. 这些性能提升也转化为了长程任务的完成能力:在固定规划器下,微调后的操作模型能够解决更多的 WebArena-Infinity 和 OpenApps 任务,其中 Qwen3.5-4B 在 119 个任务中从 31 个提升至 50 个,在 100 个任务中从 3 个提升至 15 个。这些结果表明,真实桌面环境的可控组合为提升 GUI 定位和长程计算机使用能力提供了一种可扩展的监督来源。

The framework code, the dataset, and the fine-tuned model are available from the project page: this https URL. 该框架的代码、数据集以及微调后的模型均可在项目页面获取:[点击此处访问链接]。