UI-Venus-2 Technical Report

UI-Venus-2 Technical Report

Abstract: Multimodal GUI agents have emerged as a promising paradigm for digital task automation, yet transitioning from benchmark-oriented models to dependable real-world applications remains challenging due to limited environment coverage, brittle task construction, and unreliable reward verification.

摘要: 多模态图形用户界面(GUI)智能体已成为数字任务自动化领域一种极具前景的范式。然而,由于环境覆盖范围有限、任务构建脆弱以及奖励验证不可靠,从基准测试导向的模型向可靠的现实世界应用转型仍然充满挑战。

In this work, we present UI-Venus-2, a general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework.

在这项工作中,我们提出了 UI-Venus-2,这是一个通用的基础 GUI 智能体,旨在通过统一的闭环推理-行动框架,在移动端、网页端和桌面端环境中运行。

To bridge the gap toward practical deployment, we jointly scale three critical dimensions: (1) Environments, expanding coverage to more than 170 multilingual mobile apps and native desktop operating systems; (2) Tasks, employing a deep-research pipeline for function-grounded instruction generation; and (3) Verification, adopting trace-level and sample-level evaluators with visual keypoints and multi-model voting to ensure reliable RL signals for training.

为了弥合通往实际部署的差距,我们共同扩展了三个关键维度:(1) 环境,将覆盖范围扩大到超过 170 个多语言移动应用程序和原生桌面操作系统;(2) 任务,采用深度研究流水线进行基于功能基础的指令生成;(3) 验证,采用轨迹级和样本级评估器,结合视觉关键点和多模型投票,以确保训练时强化学习(RL)信号的可靠性。

Furthermore, we integrate safety-aware mechanisms to ensure controlled execution of consequential actions. By offering a capable, efficient, and open-source foundation, UI-Venus-2 advances the field toward more generalizable, verifiable, and self-reflective agents for real-world applications.

此外,我们集成了安全感知机制,以确保对重要操作的受控执行。通过提供一个强大、高效且开源的基础,UI-Venus-2 推动了该领域的发展,使其向更具通用性、可验证性和自我反思能力的现实应用智能体迈进。


Paper Details:

  • Authors: Venus Team, et al.
  • Submission Date: 27 Aug 2026
  • Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
  • DOI: 10.48550/arXiv.2609.00028

论文详情:

  • 作者: Venus 团队等
  • 提交日期: 2026 年 8 月 27 日
  • 学科分类: 人工智能 (cs.AI);计算与语言 (cs.CL);计算机视觉与模式识别 (cs.CV);机器学习 (cs.LG)
  • DOI: 10.48550/arXiv.2609.00028