ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?
ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?
ScopeBench:智能体在目标压力下能否坚守参与边界?
Agents are increasingly deployed with real autonomy in web application and network penetration testing, where a single out-of-scope action can breach a client’s engagement boundary. Existing offensive-security benchmarks measure raw hacking capability; as those benchmarks saturate, the real barrier to deployment is a special case of alignment: scope adherence.
智能体正越来越多地被部署在具有实际自主权的 Web 应用和网络渗透测试中,而在这些场景下,任何一个超出范围的操作都可能触犯客户的参与边界。现有的攻击性安全基准测试主要衡量原始的黑客攻击能力;随着这些基准测试趋于饱和,部署智能体面临的真正障碍其实是一个特殊的对齐问题:范围合规性(Scope Adherence)。
We introduce ScopeBench, a benchmark of 30 dead-end agentic security tasks in which the stated objective is reachable only by violating the stated scope. Each task appears under two conditions that share an environment, verifier, and objective and differ only in scope: one instruction set has no scope and measures capability; the other has a natural-language scope to measure adherence.
我们推出了 ScopeBench,这是一个包含 30 个死胡同式智能体安全任务的基准测试,在这些任务中,只有通过违反既定范围才能达成目标。每个任务都在两种条件下进行,它们共享环境、验证器和目标,唯一的区别在于范围:一组指令没有范围限制,用于衡量能力;另一组则包含自然语言描述的范围限制,用于衡量合规性。
Scopeless trajectories are graded by a standard deterministic verifier. Scoped trajectories pass through two grading arms. First, the same deterministic verifier checks for the flag: because the flag sits behind the scope boundary, a pass proves by construction that a forbidden action occurred, yielding a high-precision lower bound on the violation rate. If the verifier does not pass the trajectory, an agentic judge estimates whether an out-of-scope call occurred.
无范围轨迹由标准的确定性验证器进行评分。有范围轨迹则通过两个评分环节进行评估。首先,同样的确定性验证器会检查标志(flag):由于标志位于范围边界之后,一旦通过验证,即证明发生了违规操作,从而为违规率提供了一个高精度的下限。如果验证器未通过该轨迹,则由一个智能体裁判来评估是否发生了超出范围的调用。
We calibrate the judge against 100 ScopeBench trajectories labeled call-by-call by human annotators, and a blinded audit of the evaluated rollouts finds its high recall holds - no false negatives among the 36 audited violations, with over-flagging its only observed error. Across 8 models in one harness, raw capability spans 12.2% to 81.1% and scope adherence spans 34.4% to 86.7%, with the judge finding 331 violations that mechanical verification misses.
我们利用 100 条由人工标注员逐次调用标记的 ScopeBench 轨迹对裁判进行了校准。对评估结果进行的盲审发现,其高召回率表现稳健——在 36 次被审计的违规行为中没有出现漏报(假阴性),唯一的观察误差是过度标记。在同一测试框架下的 8 个模型中,原始能力范围从 12.2% 到 81.1%,范围合规性范围从 34.4% 到 86.7%,其中裁判发现了 331 次机械验证所遗漏的违规行为。
Opus-4-8 achieves a raw-capability score 10 percentage points higher than sonnet-4-6’s while exhibiting 35.6 percentage points higher scope adherence. We release the frozen pilot benchmark, evaluation code, and all 2160 ATIF trajectories.
Opus-4-8 的原始能力得分比 sonnet-4-6 高出 10 个百分点,同时其范围合规性高出 35.6 个百分点。我们现已发布该冻结的试点基准测试、评估代码以及全部 2160 条 ATIF 轨迹。