Environment Steering: Using Data Flow Control to Improve Agent Utility and Safety

Environment Steering: Using Data Flow Control to Improve Agent Utility and Safety

环境引导:利用数据流控制提升智能体的效用与安全性

Abstract: LLM agents can make unsafe tool calls even when instructed to behave safely. Existing defenses constrain agents before execution, modify tool inputs/outputs, or rely on LLM judges; these approaches may depend on model behavior or block unsafe actions without helping the agent recover.

摘要: 大语言模型(LLM)智能体即使在被指示保持安全行为时,仍可能进行不安全的工具调用。现有的防御措施通常在执行前限制智能体、修改工具的输入/输出,或依赖 LLM 判别器;这些方法往往受限于模型行为,或者在拦截不安全操作时无法帮助智能体进行恢复。

We argue that the execution environment should instead enforce safety as the agent runs and steer it toward safe alternatives when violations occur---we call this Environment Steering.

我们认为,执行环境应当在智能体运行过程中强制执行安全性,并在发生违规时引导其转向安全的替代方案——我们将此称为“环境引导”(Environment Steering)。

We implement this by modeling the agent and harness execution state as database tables, track the record-level data flows, and check these data flows against declarative policies during runtime.

我们通过将智能体和工具的执行状态建模为数据库表来实现这一目标,追踪记录级的数据流,并在运行时根据声明式策略对这些数据流进行检查。

When violations are detected, policy- and context-specific feedback steers the agent toward safe trajectories. On AgentDyn, this enables the agent to improve task success rate over no-defense while achieving 0% attack success rate.

当检测到违规时,基于策略和上下文的反馈会引导智能体回到安全的执行轨迹。在 AgentDyn 上的测试表明,该方法在实现 0% 攻击成功率的同时,相比无防御措施的情况,显著提升了智能体的任务成功率。