Learning from Unreliable Trajectories: Adversarially-Robust Federated Q-Learning
Computer Science > Machine Learning arXiv:2610.06918 (cs) [Submitted on 2 Oct 2026] Title: Learning from Unreliable Trajectories: Adversarially-Robust Federated Q-Learning Authors: Sreejeet Maity, Aritra Mitra.
计算机科学 > 机器学习 arXiv:2610.06918 (cs) [提交于 2026 年 10 月 2 日] 标题:从不可靠轨迹中学习:对抗鲁棒的联邦 Q-学习 作者:Sreejeet Maity, Aritra Mitra。
Abstract: We study federated reinforcement learning in which multiple agents interact with a common Markov decision process and communicate through a central server to collaboratively learn the optimal state-action value function. Our goal is to understand whether the sample-efficiency benefits of collaboration can be retained when a fraction of the agents behave adversarially and transmit arbitrarily corrupted information.
摘要:我们研究了联邦强化学习,其中多个智能体与一个共同的马尔可夫决策过程交互,并通过中央服务器进行通信,以协作学习最优状态-动作价值函数。我们的目标是了解当一部分智能体表现出对抗性并传输任意损坏的信息时,协作带来的样本效率优势是否能够得以保留。
To address this problem, we introduce Robust Async-Fed-Q, an epoch-based federated learning algorithm that combines variance-reduced estimation of the Bellman optimality operator at the agents with robust aggregation at the server. We establish high-probability finite-time guarantees showing that the proposed method preserves the statistical gains of collaboration among the honest agents while tolerating adversarial corruption.
为了解决这个问题,我们引入了 Robust Async-Fed-Q,这是一种基于纪元的联邦学习算法,它将智能体端贝尔曼最优算子的方差缩减估计与服务器端的鲁棒聚合相结合。我们建立了高概率的有限时间保证,表明该方法在容忍对抗性破坏的同时,保留了诚实智能体之间协作的统计增益。
In particular, the effect of the adversarial agents decreases as the amount of data collected by each honest agent grows and eventually vanishes in the infinite-sample limit. We complement these guarantees with information-theoretic lower bounds that characterize the unavoidable statistical cost of adversarial corruption, leading to the first nearly matching upper and lower bounds for adversarially robust federated reinforcement learning.
特别地,随着每个诚实智能体收集的数据量增加,对抗性智能体的影响会减小,并最终在无限样本极限下消失。我们用表征对抗性破坏不可避免的统计成本的信息论下界补充了这些保证,从而为对抗鲁棒的联邦强化学习得出了首个几乎匹配的上下界。
We further extend our framework to accommodate single-trajectory Markovian sampling and heterogeneous partial coverage, where different agents may explore different regions of the state-action space and learning relies on their collective coverage. Finally, our epoch-based design substantially improves the best known communication complexity for federated Q-learning under asynchronous sampling.
我们进一步扩展了我们的框架,以适应单轨迹马尔可夫采样和异构部分覆盖,其中不同的智能体可能探索状态-动作空间的不同区域,学习依赖于它们的集体覆盖。最后,我们基于纪元的设计显著改善了异步采样下联邦 Q-学习的最佳已知通信复杂度。