Regularized Emphatic Temporal-Difference Learning: Stability under Constant Stepsizes
Regularized Emphatic Temporal-Difference Learning: Stability under Constant Stepsizes
正则化强调时序差分学习:常数步长下的稳定性
Abstract: Emphatic temporal-difference learning (ETD) stabilizes the expected off-policy TD update and changes its projection geometry, but neither property determines constant-stepsize sampled dynamics. We construct an ergodic two-state counterexample in which the ETD mean map contracts while the sampled product has a positive top Lyapunov exponent. Regenerative-cycle analysis separates this sign from the infinite variance of the follow-on trace.
摘要: 强调时序差分学习(ETD)稳定了预期的离策略(off-policy)TD更新并改变了其投影几何结构,但这两个特性都无法决定常数步长下的采样动力学。我们构建了一个遍历的双状态反例,其中ETD均值映射是收缩的,而采样乘积却具有正的最大李雅普诺夫指数(Lyapunov exponent)。再生循环分析将这一符号与后续迹(follow-on trace)的无限方差分离开来。
We introduce regularized emphatic TD (RETD), a normalized first-order post-shock repair that leaves the trace and importance ratios unchanged, stores the emphatic TD signal in a leaky scalar state, and releases a delayed correction. RETD’s raw equilibrium is an affine shift of the ETD equilibrium; single- and two-regularization readouts recover the ETD fixed point exactly.
我们引入了正则化强调TD(RETD),这是一种归一化的一阶冲击后修复方法,它保持迹和重要性比率不变,将强调TD信号存储在泄漏标量状态中,并释放延迟校正。RETD的原始平衡是ETD平衡的仿射变换;单正则化和双正则化读出可以精确恢复ETD的固定点。
We prove almost-sure convergence for harmonic diminishing stepsizes and a conditional constant-stepsize moment-contraction result from a Markovian random-product bound. RETD has certified negative exponents on the two-state construction and one Baird point, whereas the positive Baird ETD sign remains numerical.
我们证明了调和递减步长的几乎处处收敛性,并根据马尔可夫随机乘积界得出了条件常数步长的矩收缩结果。RETD在双状态结构和一个Baird点上被证实具有负指数,而Baird ETD的正符号目前仍仅在数值上观察到。
Paired 10,000-run experiments validate both separations, fixed-point recovery, a nonmonotone stability region, and task dependence. RETD changes post-shock dynamics; it does not reduce the shared follow-on-trace variance.
通过10,000次配对实验验证了上述分离、固定点恢复、非单调稳定性区域以及任务依赖性。RETD改变了冲击后的动力学;它并没有减少共享的后续迹方差。