Serverless gossip training of LSTM failure detectors: A matched-protocol comparison with federated, local and centralized learning on NASA C-MAPSS

Serverless gossip training of LSTM failure detectors: A matched-protocol comparison with federated, local and centralized learning on NASA C-MAPSS

基于无服务器 Gossip 协议的 LSTM 故障检测器训练:NASA C-MAPSS 数据集上联邦学习、本地学习与集中式学习的对比研究

Abstract: Industrial predictive maintenance increasingly depends on learning from equipment spread across sites whose sensor data cannot easily be pooled. Federated averaging (FedAvg) solves this with a central aggregation server; gossip learning removes the server, but its behaviour for recurrent failure-detection models has not been measured under controlled conditions.

摘要: 工业预测性维护日益依赖于对分布在不同地点的设备进行学习,而这些设备的传感器数据难以轻易汇总。联邦平均算法(FedAvg)通过中央聚合服务器解决了这一问题;Gossip 学习则去除了服务器,但其在循环故障检测模型中的表现尚未在受控条件下得到充分评估。

We compare synchronous ring gossip with FedAvg, isolated local training and a centralized reference for a stacked LSTM that detects imminent failure on the NASA C-MAPSS turbofan benchmark. All methods share one open implementation, architecture, initialization, optimizer, data split and training budget, and the primary endpoint uses one terminal window per test engine to avoid the statistical dependence of overlapping windows.

我们对比了同步环形 Gossip 协议与 FedAvg、独立本地训练以及集中式参考模型,针对 NASA C-MAPSS 涡扇发动机基准测试中的堆叠式 LSTM 故障检测模型进行了研究。所有方法均采用统一的开源实现、架构、初始化、优化器、数据划分和训练预算;主要评估指标采用每个测试引擎的单个终端窗口,以避免重叠窗口带来的统计依赖性。

On FD001 (five seeds), gossip reached a terminal-window F1 of 89.6 +/- 1.3%, compared with 89.9 +/- 1.1% for FedAvg, 83.6 +/- 6.7% for local training and 93.5 +/- 2.1% for centralized training, while transmitting the same payload as FedAvg without a coordinator. Node models agreed closely but not exactly (1.8% pairwise decision disagreement versus 5.6% without communication).

在 FD001 数据集(五个随机种子)上,Gossip 协议达到的终端窗口 F1 分数为 89.6 +/- 1.3%,相比之下,FedAvg 为 89.9 +/- 1.1%,本地训练为 83.6 +/- 6.7%,集中式训练为 93.5 +/- 2.1%。同时,Gossip 在无需协调器的情况下传输了与 FedAvg 相同的数据负载。各节点模型之间的一致性较高,但并非完全相同(两两决策分歧率为 1.8%,而无通信时为 5.6%)。

Across FD002-FD004, peer communication improved terminal-window F1 over local training by 13-28 points; gossip matched FedAvg on FD003 and FD004 but was 4.3 points lower on the multi-condition FD002 subset. Simulated message loss, node failure and server outage changed neither method appreciably, whereas larger rings degraded gossip faster. Ring gossip is therefore a practical serverless alternative when data heterogeneity is moderate, and faster-mixing topologies become important as heterogeneity grows.

在 FD002-FD004 数据集上,对等通信使终端窗口 F1 分数较本地训练提升了 13-28 个百分点;Gossip 在 FD003 和 FD004 上与 FedAvg 持平,但在多工况的 FD002 子集上低 4.3 个百分点。模拟的消息丢失、节点故障和服务器中断对两种方法的影响均不显著,但环形规模越大,Gossip 的性能下降越快。因此,当数据异构性适中时,环形 Gossip 是一种实用的无服务器替代方案;而随着异构性的增加,更快速混合(faster-mixing)的拓扑结构将变得至关重要。