
SM
Sreejeet Maity, Aritra Mitra
· 1 min read
ResearcharXiv cs.LG
Robust Federated Q-Learning with Almost No Communication
arXiv:2609.20174v1 Announce Type: new
Abstract: We consider a federated reinforcement learning setting involving $M$ agents, all of whom interact with a common Markov Decision Process (MDP). The agents exchange information via a central server to learn the optimal value function. Our goal is to understand to what extent one can hope for collaborative sample-complexity speedups in such a setting, when a small fraction of the agents are adversarial and can act arbitrarily. To that end, we propose Robust Fed-Q}, a federated Q-learning algorithm that blends ideas from both model-based and model-free RL, along with the median-of-means device from robust statistics. We prove that despite corruption, with high-probability, Robust Fed-Q (i) guarantees exact convergence to the optimal value function in the limit of infinite samples, and (ii) enjoys near-optimal finite-time rates that benefit from collaboration. In addition, our approach requires just $\tilde{O}(1)$ rounds of communication to achieve each of the above guarantees, a feature of independent interest in FL where communication is the major bottleneck.
Original source
This story was published by arXiv cs.LG and written by Sreejeet Maity, Aritra Mitra. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


