State-Space Unlearning for Non-Stationary Bias in Land Surface Forecasting

State-Space Unlearning for Non-Stationary Bias in Land Surface Forecasting

用于地表预测中非平稳偏差的状态空间“遗忘”技术

Abstract: Operational land surface forecasting systems built on Mamba-family Structured State Space Models absorb non-stationary confounding events (unrecorded irrigation booms, dam-operation shifts, sensor recalibrations) into their state-transition matrices, silently biasing NDVI, LST, and crop phenology predictions long after the physical cause ends.

摘要: 基于 Mamba 系列结构化状态空间模型(SSM)构建的业务化地表预测系统,会将非平稳的混杂事件(如未记录的灌溉激增、大坝运行调整、传感器重新校准等)吸收到其状态转移矩阵中,从而在物理原因消失后,依然长期对 NDVI(归一化植被指数)、LST(地表温度)和作物物候预测产生隐性偏差。

This paper introduces SSU-LSF (State-Space Unlearning for Land Surface Forecasting), the first machine-unlearning framework purpose-built for geoscientific Mamba-based SSMs. We develop EKFac influence functions specialized to the Mamba state matrices via a closed-form matrix-exponential gradient, use spectral-radius-weighted elbow thresholding to localize a temporal confounding footprint $\Phi$, and apply Hessian-free projected gradient ascent within a KL-divergence trust region augmented by spatial total-variation (TV) regularization.

本文提出了 SSU-LSF(地表预测的状态空间“遗忘”技术),这是首个专为地球科学领域基于 Mamba 的 SSM 模型设计的机器“遗忘”框架。我们通过闭式矩阵指数梯度,开发了专门针对 Mamba 状态矩阵的 EKFac 影响函数;利用谱半径加权的肘部阈值法来定位时间混杂足迹 $\Phi$;并在由空间全变分(TV)正则化增强的 KL 散度信任域内,应用了无 Hessian 矩阵的投影梯度上升法。

Proposition 1 establishes that residual confounding is bounded by $\mathcal{O}\big((1-\rho(\bar{A})^{T_c})/((1-\rho(\bar{A}))\mu)\big)$, which grows with the window length $T_c$. Across three heterogeneous benchmarks and eleven baselines, SSU-LSF achieves confounding reduction rates of $0.773$ (CropHarvest), $0.821$ (NDVI-LST), and $0.859$ (ERA5), with worst-case clean-domain RMSE degradation of $4.2%$ on ERA5, converging in 3—5 epochs at $8.4\times$ lower GPU-cost per unlearning request than full retraining.

命题 1 证明了残余混杂项受限于 $\mathcal{O}\big((1-\rho(\bar{A})^{T_c})/((1-\rho(\bar{A}))\mu)\big)$,该值随窗口长度 $T_c$ 的增加而增大。在三个异构基准测试和十一个基线模型中,SSU-LSF 在 CropHarvest、NDVI-LST 和 ERA5 数据集上的混杂消除率分别达到了 $0.773$、$0.821$ 和 $0.859$;在 ERA5 上的最差清洁域 RMSE 降幅仅为 $4.2%$,且在 3—5 个 epoch 内即可收敛,单次“遗忘”请求的 GPU 成本比完全重新训练降低了 $8.4$ 倍。