How Far is Adam from Natural Gradient Descent?
How Far is Adam from Natural Gradient Descent?
Adam 距离自然梯度下降(NGD)有多远?
Abstract: Adam is the standard optimizer in deep learning, yet its geometric relationship to natural gradient descent (NGD) contains unresolved questions. We study Adam’s full update rule, including momentum, as a diagonal empirical Fisher approximation subject to diagonal truncation, empirical label substitution, and temporal lag.
摘要: Adam 是深度学习中的标准优化器,但它与自然梯度下降(NGD)之间的几何关系仍存在未解之谜。我们将 Adam 的完整更新规则(包括动量)视为一种对角经验 Fisher 近似,并对其进行了对角截断、经验标签替换和时间滞后处理。
Using the scale-invariant $\gamma(\Delta\theta)$ metric, we measure Adam’s geometric deviation from true NGD across four loss landscapes: well-conditioned linear regression, ill-conditioned linear regression, logistic regression, and a non-convex small neural network. Adam’s geometric trajectory is context-dependent. Deviation remains low in well-conditioned settings but rises significantly under ill-conditioning, reaching misalignments of $\approx 10^3$ in the neural network.
我们使用尺度不变的 $\gamma(\Delta\theta)$ 度量,在四种损失函数场景下测量了 Adam 与真实 NGD 之间的几何偏差:良态线性回归、病态线性回归、逻辑回归以及一个非凸的小型神经网络。Adam 的几何轨迹具有上下文依赖性。在良态设置下,偏差保持在较低水平,但在病态条件下偏差显著上升,在神经网络中失准度达到了 $\approx 10^3$。
Higher geometric drift correlates with slower initial optimization but does not degrade final objective minimization; Adam consistently reaches low loss. Furthermore, the improved empirical Fisher (iEF) tracks more stable paths than the standard empirical Fisher (EF), which frequently oscillates or diverges. Our results suggest Adam’s practical optimization power may stem from a balance of structural approximation errors and momentum smoothing rather than close tracking of the natural gradient path.
较高的几何漂移与较慢的初始优化速度相关,但并不会降低最终的目标函数最小化效果;Adam 始终能达到较低的损失值。此外,改进后的经验 Fisher(iEF)比标准的经验 Fisher(EF)能追踪到更稳定的路径,后者经常出现震荡或发散。我们的研究结果表明,Adam 的实际优化能力可能源于结构近似误差与动量平滑之间的平衡,而非对自然梯度路径的紧密追踪。