Learning When to Refine: Long-Horizon Reinforcement Learning for Budgeted Neural-Operator PDE Solvers

Computer Science > Machine Learning arXiv:2610.06883 (cs) [Submitted on 19 Sep 2026] Title: Learning When to Refine: Long-Horizon Reinforcement Learning for Budgeted Neural-Operator PDE Solvers Authors: Ange Tong.

计算机科学 > 机器学习 arXiv:2610.06883 (cs) [提交于 2026 年 9 月 19 日] 标题:学习何时细化:用于预算神经算子偏微分方程求解器的长程强化学习 作者:Ange Tong。

Abstract: Neural operators provide fast surrogates for time-dependent PDEs, but autoregressive deployment creates a refinement-allocation problem: prediction errors vary over space and time, while only a finite number of local corrections can be committed along a trajectory. We formulate this as budgeted adaptive neural-operator solving.

摘要:神经算子为随时间变化的偏微分方程提供了快速代理模型,但自回归部署产生了一个细化分配问题:预测误差在空间和时间上各不相同,而沿轨迹只能进行有限次数的局部修正。我们将此问题表述为预算自适应神经算子求解。

A global Fourier neural operator advances the full field, a local operator proposes patch-wise residual corrections, and a set-aware selector chooses where to refine. A macro policy decides when and how much of the remaining refinement budget to spend.

全局傅里叶神经算子推进整个场,局部算子提出分块残差修正,集合感知选择器决定细化位置。宏观策略决定何时以及消耗多少剩余的细化预算。

We introduce rollout-verified policy improvement (RV-PI), which evaluates feasible refinement counts through actual continuation rollouts of the learned PDE solver, converts long-horizon advantages into conservative policy targets, and accepts an update only when held-out trajectory error improves.

我们引入了滚动验证策略改进(RV-PI),它通过学习到的偏微分方程求解器的实际连续滚动来评估可行的细化次数,将长程优势转化为保守的策略目标,并仅在留出轨迹误差改善时才接受更新。

On the shallow-water benchmark with a 32-intervention budget, RV-PI achieves a three-seed mean trajectory relative L2 error of 0.6910, improving over immediate-only policy improvement by 5.37% and RandomMacro by 2.41%.

在具有 32 次干预预算的浅水基准测试中,RV-PI 实现了 0.6910 的三种子平均轨迹相对 L2 误差,比仅即时策略改进提高了 5.37%,比 RandomMacro 提高了 2.41%。

On the forcing-driven Brusselator benchmark with a 76-intervention budget, RV-PI attains 0.09954, improving over immediate-only policy improvement by 2.31% and RandomMacro by 5.32%. These results show that, under a fixed refinement budget, the value of a local correction depends on its downstream effect on the autoregressive trajectory, not only on its immediate error reduction.

在具有 76 次干预预算的强迫驱动 Brusselator 基准测试中,RV-PI 达到了 0.09954,比仅即时策略改进提高了 2.31%,比 RandomMacro 提高了 5.32%。这些结果表明,在固定的细化预算下,局部修正的价值取决于其对自回归轨迹的下游影响,而不仅仅取决于其即时的误差减少。