DeReAct: Decomposed Reasoning and Acting for Reliable AI Agents

DeReAct: Decomposed Reasoning and Acting for Reliable AI Agents

DeReAct:用于可靠 AI 智能体的分解推理与行动框架

Abstract: ReAct-based agents typically rely on a single LLM policy to propose actions, interact with the environment, and decide when a task is complete. This coupling makes action authorization and completion control difficult to enforce independently, allowing errors to propagate and unsupported completion claims to terminate execution.

摘要: 基于 ReAct 的智能体通常依赖单一的大语言模型(LLM)策略来提出行动建议、与环境交互并决定任务何时完成。这种耦合使得行动授权和完成控制难以独立执行,从而导致错误传播,并可能因未经证实的完成声明而导致执行过早终止。

We introduce DeReAct, a modular agent architecture that externalizes two gating policies: a Critic that validates proposed actions before execution, and a Context Manager that reconstructs an environment-supported \textsc{State} and certifies task completion.

我们引入了 DeReAct,这是一种模块化的智能体架构,它将两个门控策略外部化:一个是负责在执行前验证建议行动的“评论家”(Critic),另一个是负责重构环境支持的“状态”(State)并认证任务完成情况的“上下文管理器”(Context Manager)。

Across GAIA and SWE-bench Verified, DeReAct improves Pass@1 most for weaker Brain models, with gains of 6.5—7.0 points for Qwen3-Coder-480B and 4.2—5.2 points for Claude Sonnet 4.5; gains diminish as Brain capability increases.

在 GAIA 和 SWE-bench Verified 测试集上,DeReAct 对较弱的“大脑”模型(Brain models)的 Pass@1 指标提升最为显著,其中 Qwen3-Coder-480B 提升了 6.5–7.0 个百分点,Claude Sonnet 4.5 提升了 4.2–5.2 个百分点;随着“大脑”模型能力的增强,这种增益会逐渐减小。

Trajectory and ablation analyses show that external gating is effective when targeted failures are sufficiently prevalent and the gating policy is itself sufficient. With Claude Opus 4.5, Pass@1 remains comparable to ReAct, while DeReAct produces more evidence-complete and constraint-satisfying trajectories, indicating that completion control can trade earlier termination for stronger grounding.

轨迹和消融分析表明,当目标故障足够普遍且门控策略本身足够有效时,外部门控是有效的。在使用 Claude Opus 4.5 时,Pass@1 指标与 ReAct 相当,但 DeReAct 生成了证据更完整且更符合约束条件的轨迹,这表明完成控制可以通过牺牲更早的终止时间来换取更强的基础支撑(Grounding)。

Overall, DeReAct improves weaker agents while retaining grounding benefits as models strengthen.

总的来说,DeReAct 在提升较弱智能体性能的同时,在模型增强时依然保留了基础支撑方面的优势。