Response Renormalization for Critical Deep Equilibrium Models
Response Renormalization for Critical Deep Equilibrium Models
临界深度平衡模型的响应重归一化
Abstract: Deep Equilibrium Models (DEQs) compute predictions from a hidden representation unchanged by the model update. Training through this equilibrium uses implicit differentiation and requires solving an adjoint system built from the residual Jacobian. If this Jacobian is nearly singular along loss-sensitive directions, small perturbations can be strongly amplified in the adjoint response, producing large, highly sensitive gradients that can make optimization unreliable.
摘要: 深度平衡模型(DEQs)通过模型更新后保持不变的隐藏表示来计算预测。通过这种平衡进行训练需要使用隐式微分,并求解由残差雅可比矩阵构建的伴随系统。如果该雅可比矩阵在损失敏感方向上接近奇异,微小的扰动可能会在伴随响应中被强烈放大,从而产生巨大且高度敏感的梯度,导致优化过程变得不可靠。
We introduce Response Renormalization, a backward-pass framework that lifts selected near-pole denominators while leaving unlifted response channels unchanged. Collective Mode Response Renormalization (CMR) applies this correction in a low-dimensional critical subspace, while Phi-adaptive CMR computes a bounded response mass from a positive susceptibility rule.
我们引入了“响应重归一化”(Response Renormalization),这是一种反向传播框架,它能够提升选定的近极点分母,同时保持未提升的响应通道不变。集体模式响应重归一化(CMR)在低维临界子空间中应用此校正,而 Phi 自适应 CMR 则通过正敏感度规则计算有界的响应质量。
We derive dense and matrix-free collective formulations, distinguish exact gradients of a modified frozen-anchor residual from backward-response surrogates, and extend the construction to Structured Implicit Layers and Vector Attractors (SILVA).
我们推导了稠密和无矩阵的集体公式,区分了修改后的固定锚点残差的精确梯度与反向响应代理,并将该结构扩展到结构化隐式层和向量吸引子(SILVA)。
Across 23 multiphysics families spanning partial differential equations, three-dimensional fields, operator maps, complex geometries, and particle systems, CMR and Phi-CMR yield test errors no more than five percent higher than those from models trained with exact implicit differentiation in more than 98% of static and 95% of transient family-seed comparisons.
在涵盖偏微分方程、三维场、算子映射、复杂几何结构和粒子系统的 23 个多物理场系列中,在超过 98% 的静态和 95% 的瞬态系列种子比较中,CMR 和 Phi-CMR 的测试误差比使用精确隐式微分训练的模型高出不超过 5%。
Solver-index experiments show convergence toward the static adjoint, while physical-time rollouts retain predictive fidelity under the evaluated conditions. These results demonstrate that selective response renormalization can control near-critical adjoint amplification without globally damping well-conditioned sensitivity. Therefore, the method can make parameter updates more reliable while preserving the useful gradient information needed for learning.
求解器索引实验显示了向静态伴随收敛的趋势,而物理时间展开在评估条件下保持了预测保真度。这些结果表明,选择性响应重归一化可以在不全局抑制良性敏感度的情况下,控制近临界伴随放大。因此,该方法能够在保留学习所需有用梯度信息的同时,使参数更新更加可靠。