Transferability and operational reliability of a Prithvi crop classification foundation model under phenological and geographic shift across three continents
Fine-tuned geospatial foundation models (GeoFMs) pretrained on large satellite archives have been shown to improve crop classification accuracy and geographic transferability. However, their operational performance beyond the training distribution remains poorly characterized. We evaluated the out-of-distribution performance of a widely adopted GeoFM [Prithvi-EO-2.0] across 37 events in 12 countries on three continents and validated against regional reference products.
在大型卫星档案上进行预训练的微调地理空间基础模型(GeoFMs)已被证明可以提高作物分类的准确性和地理迁移能力。然而,它们在训练分布之外的操作性能仍然缺乏充分的表征。我们评估了一种广泛采用的 GeoFM [Prithvi-EO-2.0] 在三大洲 12 个国家的 37 个事件中的分布外性能,并针对区域参考产品进行了验证。
Results indicated that the mean overall accuracy (OA) declined from 0.65 in the United States to 0.40 in Europe. Beyond accuracy metrics, we assessed five key aspects of model performance: whether model confidence indicates signal failure, sensitivity to observation windows, the effect of coarsening class schemes, and robustness to both band loss and cloud- and shadow-contamination.
结果表明,平均总体准确率(OA)从美国的 0.65 下降到欧洲的 0.40。除了准确性指标外,我们还评估了模型性能的五个关键方面:模型置信度是否能指示信号故障、对观测窗口的敏感性、粗化分类方案的影响,以及对波段丢失和云影污染的鲁棒性。
Accuracy collapsed when the observation window misaligned with local crop phenology, while deterministic confidence remained high. Expected calibration error increased for seven of eight paired events, and 12-51% of each affected scene was confidently mislabeled at near-zero precision. Monte Carlo dropout entropy registered the shift in all eight, indicating that much of the apparent cross-continent decline reflected phenological misalignment rather than spatial transfer.
当观测窗口与当地作物物候不匹配时,准确率会崩溃,而确定性置信度却保持在高位。在八个配对事件中的七个中,预期校准误差增加,并且每个受影响场景中有 12-51% 的部分在近乎零精度的情况下被自信地错误标记。蒙特卡洛 dropout 熵记录了所有八个事件中的这种偏移,表明大部分明显的跨洲下降反映的是物候不匹配,而非空间迁移问题。
Two adjustments recovered accuracy without retraining. Consolidating 13 classes into 10, based on the model’s dominant confusions, raised the mean OA by 8.4 percentage points. Compressing the window toward near-real-time use preserved accuracy across a 45- to 90-day plateau, peaking near 75 days, though arms tighter than 30 days fell about 0.11 below that plateau. Fine-tuned crop GeoFMs therefore transfer usefully only where observation windows match local growing seasons.
无需重新训练,通过两项调整即可恢复准确性。基于模型的主要混淆情况,将 13 个类别合并为 10 个,使平均 OA 提高了 8.4 个百分点。将窗口压缩至接近实时使用,在 45 到 90 天的平稳期内保持了准确性,并在 75 天左右达到峰值,尽管小于 30 天的窗口期比该平稳期低约 0.11。因此,微调后的作物 GeoFMs 只有在观测窗口与当地生长季节匹配时才能有效迁移。