Taming Staleness in Asynchronous Federated Learning With System Heterogeneity
Abstract
Asynchronous federated learning is a distributed machine learning framework designed for parallel and distributed systems. However, system heterogeneity in communication and computation can cause stragglers to upload stale updates, which may degrade the performance of the global model. Existing studies primarily focus on systems with mild to moderate heterogeneity, leaving the trade-off between communication efficiency and model performance insufficiently addressed in extreme scenarios. To tackle this issue, we propose FedASRA, an efficient semi-asynchronous federated learning framework featuring adaptive alternating selection (AAS) and real-time staleness adjustment (RSA). Specifically, AAS leverages a multi-attribute utility function to periodically schedule the participation of stragglers and fast clients in an alternating manner, thereby narrowing the system heterogeneity gap among clients in each communication round. When staleness is unavoidable, RSA enables stragglers to synchronously receive the latest global model during training and compute updated gradients. These updated gradients are then locally aggregated with the soon-to-be stale gradients to improve the freshness of submitted updates. We provide convergence and system complexity analysis for FedASRA, establishing a staleness-agnostic convergence upper bound. Extensive experiments are conducted in simulated environments and a prototype system built with 21 real-world devices. Results demonstrate that FedASRA achieves a favorable balance between system efficiency and performance. Compared with state-of-the-art methods, FedASRA reduces wall-clock time by 18.51% and communication overhead by 11.66% on average.