A Reinforcement Learning–Driven Latency Optimization Framework for Heterogeneous Federated Learning on Edge Devices
Abstract
Heterogeneous federated learning (FL) over edge networks suffers from high end-to-end latency due to coupled delays in model distribution, on-device training and upload, and server-side aggregation. Existing latency-aware FL methods typically optimize only a single stage, such as client scheduling or communication compression, and therefore remain suboptimal under device heterogeneity. This paper proposes a server-centric orchestration framework that jointly models and optimizes the entire FL pipeline through a three-stage latency model, multi-factor client selection, hierarchical aggregation, and reinforcement-learning-driven decision making. The proposed framework treats heterogeneous FL orchestration as a coupled optimization problem rather than an isolated client-selection task, jointly coordinating participant selection, hierarchical aggregation, and latency-aware server-side decisions through a unified three-stage latency model. The server’s sequential orchestration is formulated as a Markov decision process and optimized via Proximal Policy Optimization. Theoretical discussion provides intuition on the compatibility of the proposed design with commonly used FL convergence assumptions under bounded heterogeneity, rather than establishing a formal convergence guarantee. Multi-seed simulation experiments across five network scales ranging from 5 to 50 devices show that the proposed framework maintains competitive end-to-end latency relative to strong baselines, while reducing average communication latency compared with non-learned selection in the main policy-comparison experiments. The main empirical value of the framework lies in demonstrating that jointly coupled server-side orchestration can absorb its own control overhead and remain effective under heterogeneous FL conditions.