AutoSynthData: Generating Training Data for Enterprise Agents

AutoSynthData: Generating Training Data for Enterprise Agents

AutoSynthData:为企业智能体生成训练数据

Enterprises need agents that work well in their own environments. The work they ask these agents to do is shaped by the systems they use, the rules they follow, and the state of their data. A model may be broadly capable and still struggle with a particular environment: a workflow it handles poorly, a combination of tools it misuses, or a constraint it fails to respect. Those are the weaknesses an enterprise needs to improve. 企业需要能够在自身环境中良好运行的智能体。企业要求这些智能体完成的工作,受到其所用系统、所遵循规则以及数据状态的制约。一个模型可能具备广泛的能力,但在特定环境中仍可能表现不佳:例如处理不当的工作流、误用的工具组合,或是未能遵守的约束条件。这些正是企业需要改进的弱点。

The difficulty is turning those weaknesses into training data. An individual failure tells us something, but training a model requires many new tasks that exercise the same capability in different situations. Those tasks must also be possible to complete in the environment, resemble work someone would actually request, and have a reliable way to check whether the agent succeeded. 难点在于如何将这些弱点转化为训练数据。单次的失败能提供一些信息,但训练模型需要大量新的任务,在不同情境下反复锻炼同一能力。这些任务还必须满足以下条件:在环境中可执行、类似于用户实际会提出的需求,并且拥有可靠的方法来验证智能体是否成功。

At ServiceNow CoreAI, we built AutoSynthData to turn those capability gaps into training data. It uses a target model’s failures and a stronger teacher’s successes to decide what the model should learn next, then generates and validates new tasks that exercise those capabilities. As the model improves, the curriculum shifts toward what it still finds difficult. We illustrate the pipeline with EnterpriseOps Gym (Malay et al., 2026), using the released dataset. 在 ServiceNow CoreAI,我们构建了 AutoSynthData,旨在将这些能力差距转化为训练数据。它利用目标模型的失败案例和更强教师模型的成功案例,来决定模型下一步的学习重点,随后生成并验证能够锻炼这些能力的新任务。随着模型的改进,课程内容会转向模型仍感到困难的部分。我们通过 EnterpriseOps Gym (Malay et al., 2026) 并使用已发布的数据集,展示了这一流程。

We begin by describing the environment an agent operates in and what makes a task useful for training. What makes a useful agentic task? An agentic environment defines the world in which an agent operates: the state it can observe and modify, the tools and APIs it can invoke, and the state transitions produced by its actions. A task is instantiated within this environment. We use the following abstraction: task = (system specification, user prompt, verifier) 我们首先描述智能体运行的环境,以及什么样的任务对训练有价值。什么是有用的智能体任务?智能体环境定义了智能体运行的世界:它能观察和修改的状态、它能调用的工具和 API,以及由其行为产生的状态转换。任务就在此环境中实例化。我们使用以下抽象:任务 = (系统规范, 用户提示, 验证器)。

System specification: The system specification defines the constraints under which the agent operates, including system instructions, environment policies, and, when applicable, task-specific initialization such as a seeded database state or a set of knowledge articles. The specification must be compatible with the environment’s tools, state, and supported actions. Its instructions should be clear and avoid arbitrary constraints introduced solely to manufacture difficulty. 系统规范:系统规范定义了智能体运行的约束条件,包括系统指令、环境策略,以及在适用时针对特定任务的初始化设置(如预置的数据库状态或知识库文章)。该规范必须与环境的工具、状态和支持的操作兼容。其指令应清晰明确,避免为了制造难度而引入任意的约束。

Agent-facing task: The user prompt specifies what the user wants the agent to accomplish, together with any user-level constraints. A generated task should satisfy three properties: Feasibility (there should exist at least one trajectory that satisfies the user prompt while respecting the system specification), Realism (the user prompt should resemble something a user would plausibly ask), and Difficulty (the task should expose a weakness of the current agent). 面向智能体的任务:用户提示指定了用户希望智能体完成的目标,以及任何用户层面的约束。生成的任务应满足三个属性:可行性(必须存在至少一条既满足用户提示又符合系统规范的路径)、真实性(用户提示应类似于用户在目标环境中可能提出的需求)、难度(任务应能暴露当前智能体的弱点)。

Verifier: The verifier determines whether the resulting trajectory successfully completes the task. It should satisfy three properties: Consistency (agreeing with the prompt, specification, and state), Soundness (rejecting trajectories that fail or violate constraints), and Completeness (accepting valid solutions rather than encoding one particular reference trajectory). 验证器:验证器用于判断生成的轨迹是否成功完成了任务。它应满足三个属性:一致性(与提示、规范和环境状态保持一致)、稳健性(拒绝未能完成任务或违反约束的轨迹)、完备性(接受所有有效的解决方案,而不是仅限于某一条特定的参考轨迹)。

Overview: Given an environment and a target model, AutoSynthData generates training tasks consisting of a system specification, user prompt, and verifier. The generated tasks are grounded in the environment and selected to provide useful training signal for the current model. AutoSynthData first evaluates the target model in the environment using diagnostic tasks and identifies patterns in the tasks it struggles to complete. A stronger teacher helps characterize which of those tasks are solvable and what successful behavior looks like. 概述:给定一个环境和一个目标模型,AutoSynthData 会生成包含系统规范、用户提示和验证器的训练任务。生成的任务基于环境构建,并经过筛选,旨在为当前模型提供有效的训练信号。AutoSynthData 首先使用诊断任务在环境中评估目标模型,并识别其难以完成的任务模式。更强的教师模型有助于界定哪些任务是可解的,以及成功的行为表现是什么样的。

From model failures to a curriculum: AutoSynthData uses evaluation runs in the target environment to identify what the model needs to learn next. In our EnterpriseOps Gym experiment, we run both the target model and a stronger teacher on the evaluation tasks. We examine those runs to identify: the capability being tested; the tools and workflow structure involved; where the target model fails and how the teacher succeeds; the properties that a correct final state must satisfy; the dimensions that can vary while preserving the capability being tested. 从模型失败到课程体系:AutoSynthData 利用目标环境中的评估运行来确定模型下一步需要学习的内容。在我们的 EnterpriseOps Gym 实验中,我们让目标模型和更强的教师模型同时执行评估任务。我们检查这些运行结果以确定:正在测试的能力;涉及的工具和工作流结构;目标模型在何处失败以及教师模型如何成功;正确最终状态必须满足的属性;以及在保持测试能力的前提下可以变化的维度。

Generating and scaling tasks: Identifying a capability gap tells us what to teach, but training requires many varied tasks that exercise it. AutoSynthData uses the specification card to generate… 生成与扩展任务:识别出能力差距告诉了我们教什么,但训练需要大量多样的任务来反复练习。AutoSynthData 使用规范卡片来生成……