Welcome RL Environments to the hub
Welcome RL Environments to the hub
Reinforcement Learning environments give new capabilities to agentic AI systems, and they’re a great way to measure and improve performance in your agents. Therefore, Hugging Face Hub now has a special place for RL Environments. 强化学习(RL)环境为智能体 AI 系统赋予了新的能力,也是衡量和提升智能体性能的绝佳方式。因此,Hugging Face Hub 现在为强化学习环境开辟了一个专门的板块。
An environment gives an agent a task, responds to its actions with observations, and scores the outcome. The resulting rewards can measure an agent’s performance during evaluation or provide a learning signal during training. For an introduction to this interaction loop, see our blogpost on environments. 环境为智能体提供任务,通过观测值响应其动作,并对结果进行评分。由此产生的奖励既可以在评估期间衡量智能体的表现,也可以在训练期间提供学习信号。有关此交互循环的介绍,请参阅我们关于环境的博客文章。
Within the environment, the agent will perform a set of tasks that are represented as datasets. Therefore, environments can be split into broadly two parts: tasksets and runtimes. In this release, we are focusing on the tasksets. 在环境中,智能体将执行一系列以数据集形式呈现的任务。因此,环境大致可以分为两部分:任务集(tasksets)和运行时(runtimes)。在本次发布中,我们重点关注任务集。
An RL environment on the Hub is a dataset repo that shows up in the new RL Environments filter. The “Use this dataset” button gives you the command to run it in that framework. There is no new repo type, no registry, and no sign-up. There are already environments in Harbor, Verifiers, and NVIDIA NeMo Gym. Hub 上的强化学习环境是一个数据集仓库,会显示在新的“RL Environments”筛选器中。“Use this dataset”按钮会提供在相应框架中运行该环境的命令。这里没有新的仓库类型,无需注册,也无需额外申请。目前,Harbor、Verifiers 和 NVIDIA NeMo Gym 中已经包含了相关环境。
Stop building environment registries
停止构建环境注册表
Every RL paper or framework uses its own way to find environments. Custom hubs, runtime registries, independent task datasets, or a GitHub list of tasks with a custom loader. This means that many of the published environments are siloed: if you publish an environment for one framework, users of the other three can’t load it. If you want to train on an environment from another framework or a new paper, you’ll need to port it by hand. 每一篇强化学习论文或框架都有自己寻找环境的方式。无论是自定义中心、运行时注册表、独立任务数据集,还是带有自定义加载器的 GitHub 任务列表,这意味着许多已发布的环境都是孤立的:如果你为一个框架发布了环境,其他三个框架的用户就无法加载它。如果你想在另一个框架或新论文的环境中进行训练,就必须手动移植它。
We think this is the wrong shape. An environment is tasks, tests, containers, and a reward rule, which are data with a runtime on top. The Hub already stores data, versions it, gates it, previews it, and serves it to millions of people. It does not need a second system to hold environments. It needs a way to say “this data is an environment, and here is how you run it.” 我们认为这种模式是不对的。环境本质上是任务、测试、容器和奖励规则的集合,即“数据 + 运行时”。Hub 已经具备了存储、版本控制、权限管理、预览数据并将其服务于数百万人的能力。它不需要第二个系统来托管环境,它只需要一种方式来声明:“这些数据是一个环境,而这是运行它的方法。”
Task data lives on the Hub. Runtime configuration and verifier code can live in the repo or the framework. The frameworks keep doing what they are good at. The Hub does what it is good at, which is hosting, discovery, and versioning. Nobody has to own the catalogue. In fact, catalogues can run on other platforms too, powered by the hub. 任务数据存储在 Hub 上。运行时配置和验证器代码可以存在于仓库或框架中。框架继续做它们擅长的事,而 Hub 则做它擅长的事,即托管、发现和版本控制。没有人需要垄断目录。事实上,目录也可以在其他平台上运行,并由 Hub 提供支持。
The dataset repository hosts your environment files. The framework runs them locally or on a supported cloud backend. Hugging Face Jobs can run cloud workloads, and Hugging Face Sandboxes, built on Jobs, provide interactive command execution. The tags describe compatibility and generate loading commands; adding a tag does not start a job or sandbox. 数据集仓库托管你的环境文件。框架在本地或受支持的云后端上运行它们。Hugging Face Jobs 可以运行云工作负载,而基于 Jobs 构建的 Hugging Face Sandboxes 则提供交互式命令执行。标签用于描述兼容性并生成加载命令;添加标签本身不会启动任务或沙盒。
What shipped
本次发布内容
The RL Environments filter. Go to huggingface.co/datasets?other=rl-environment. Every dataset with the rl-environment tag appears there, whatever framework it works with.
RL Environments 筛选器。 前往 huggingface.co/datasets?other=rl-environment。无论使用何种框架,所有带有 rl-environment 标签的数据集都会出现在那里。
Framework tags. Four environment frameworks are registered as dataset libraries: 框架标签。 四个环境框架已注册为数据集库:
| Tag | Framework |
|---|---|
| harbor | Harbor |
| verifiers | Verifiers |
| openenv | OpenEnv |
| nemo-gym | NeMo Gym |
Each framework tag puts the framework’s icon on the dataset page and adds a generated snippet to “Use this dataset”. A dataset can carry more than one framework tag. That is the point. Tags describe compatibility, and compatibility is not exclusive. Each listed framework must support the files in the repository; adding a tag does not convert them. 每个框架标签都会在数据集页面上显示该框架的图标,并向“Use this dataset”添加生成的代码片段。一个数据集可以携带多个框架标签,这正是设计的初衷。标签描述的是兼容性,而兼容性并非排他的。每个列出的框架都必须支持仓库中的文件;添加标签并不会自动转换文件格式。
Run an environment and inspect its reward
运行环境并检查奖励
Choose the example for your framework and run it in a separate Python environment with the prerequisites listed below. 选择适合你框架的示例,并在下方列出的先决条件下,在独立的 Python 环境中运行它。
Harbor: run a reference solution Harbor:运行参考解决方案
Harbor can load task directories from a Hub repository. The oracle agent runs the task’s reference solution, then the verifier scores the result. It does not call a model. Harbor 可以从 Hub 仓库加载任务目录。Oracle 智能体会运行任务的参考解决方案,然后由验证器对结果进行评分。它不会调用模型。
uv tool install --python 3.13 'harbor==0.21.0'
harbor run \
--repo https://huggingface.co/datasets/harborframework/terminal-bench-2.1 \
--dataset terminal-bench-2.1@2.1.0 \
--include-task-name '*regex-log' \
--agent oracle --env docker --jobs-dir results/harbor
harbor view results/harbor
The viewer shows the task’s reward, verifier output, and logs. This checks the task and its reference solution before you try a model agent. 查看器会显示任务的奖励、验证器输出和日志。在尝试模型智能体之前,这可以检查任务及其参考解决方案是否正常。
Verifiers: run a model on the same task Verifiers:在同一任务上运行模型
The Harbor integration of verifiers v1 can run the same task directories in different runtimes, such as Docker. It also supports different harnesses, including a minimal bash harness. Verifiers v1 的 Harbor 集成可以在不同的运行时(如 Docker)中运行相同的任务目录。它还支持不同的测试工具(harnesses),包括一个极简的 bash 测试工具。
uvx --python 3.13 --from 'verifiers[harbor]' eval harbor \
--env.taskset.repo https://huggingface.co/datasets/harborframework/terminal-bench-2.1 \
--env.taskset.dataset terminal-bench-2.1@2.1.0 \
--env.taskset.tasks '["regex-log"]' \
--env.agent.runtime.type docker \
--env.agent.harness.id bash \
--model "$MODEL" \
--client.base-url "$LLM_URL"
Here repo is the full Hugging Face Git URL, while dataset is the name and version in that repo’s registry.json. This loader uses Harbor’s registry conventions, so a bare Hub repo ID cannot replace both values.
这里的 repo 是完整的 Hugging Face Git URL,而 dataset 是该仓库 registry.json 中的名称和版本。此加载器使用 Harbor 的注册表约定,因此仅凭 Hub 仓库 ID 无法替代这两个值。
OpenEnv: run an agent and inspect its reward OpenEnv:运行智能体并检查奖励
OpenEnv’s Harbor integration can run the same task directories with an agent such as OpenCode and return the verifier’s reward alongside the agent’s trace. OpenEnv 的 Harbor 集成可以使用 OpenCode 等智能体运行相同的任务目录,并返回验证器的奖励以及智能体的执行轨迹。
pip install "openenv[harbor]==0.7.0"
openenv harbor rollout \
--llm-url "$LLM_URL" \
--model "$MODEL" \
--dataset harborframework/terminal-bench-2.1 \
--task-index 0 \
--harness opencode \
--sandbox docker \
--out rollout.json
The command downloads the dataset’s tasks/ directories, runs one task in Docker, and writes the result. The default connection uses a temporary Gradio tunnel so the sandboxed agent can reach OpenEnv’s model proxy.
该命令会下载数据集的 tasks/ 目录,在 Docker 中运行一个任务,并写入结果。默认连接使用临时的 Gradio 隧道,以便沙盒中的智能体能够访问 OpenEnv 的模型代理。
Read the verifier result and the number of model calls: 读取验证器结果和模型调用次数:
import json
from pathlib import Path
result = json.loads(Path("rollout.json").read_text())[0]
print("Re