HybridFLow is presented, a closed-loop SDN-driven orchestration framework that integrates network-layer intelligence directly into hybrid FL and generates calibrated per-client communication-time estimates before each training round and uses them to partition clients into synchronous and asynchronous groups while balancing round latency and update staleness.
Abstract
Cross-silo Federated Learning (FL) enables geographically distributed institutions to collaboratively train machine learning models without sharing raw data. In wide-area deployments, however, communication delays often dominate round completion time and exacerbate the straggler effect. Hybrid FL addresses this challenge by combining synchronous and asynchronous client participation, but effective partitioning requires visibility into network conditions such as shared bottlenecks, link utilization, and path contention that individual clients cannot observe. We present HybridFLow, a closed-loop SDN-driven orchestration framework that integrates network-layer intelligence directly into hybrid FL. Leveraging the SDN controller's global topology view, HybridFLow generates calibrated per-client communication-time estimates before each training round and uses them to partition clients into synchronous and asynchronous groups while balancing round latency and update staleness. After each round, measured communication times are fed back to the controller to continuously refine future predictions. Experimental results across multiple network topologies show that HybridFLow reaches 80% target accuracy 33-40% faster than SmartFLow and reduces average round duration by 30-40 seconds, while FedAsync fails to reach the target accuracy under non-IID data distributions.
Heterogeneous federated learning (FL) over edge networks suffers from high end-to-end latency due to coupled delays in model distribution, on-device training and upload, and server-side aggregation. Existing latency-aware FL methods typically optimize only a single stage, such as client scheduling or communication comp...
Split Federated Learning (SFL) has emerged as a pivotal paradigm for privacy-preserving distributed training on resource-constrained edge devices by partitioning neural networks between clients and a server. A critical design choice in SFL is the split layer, which determines the computation distribution and the semant...
Ai-Jing Li, Ya-Wen Li, Guan-Hua Ye et al.· Proceedings of the Thirty-Fi...· 0 citations
An asynchronous framework named FedQS, which employs a multi-dimensional staleness evaluation mechanism that dynamically assesses updates by combining the similarity between local and global models with client latency metrics, and implements a decoupling solution via a queue scheduling algorithm to resolve the coupling...
Jia-Hui Zhou, Fang Li, Tian-Yu Shi et al.· Journal of Cloud Computing· 0 citations
An SDN-orchestrated architecture for delay control in wide-area networks, combining edge-based GRU Active Queue Management (AQM) with layer-wise federated averaging, indicates that federated averaging reduces rare congestion events without meaningful degradation of normal bottleneck operation.
Karol Marszałek, Adam Domański· Scientific Reports· 0 citations
Federated learning (FL) is increasingly deployed as a managed learning service rather than as a set of isolated training jobs. In networked edge environments, dependent FL service flows must coordinate heterogeneous clients, non-IID data, fluctuating communication latency, and precedence-constrained tasks under service...
Jie-Ping Luo, Qi-Yue Li, Yuxuan Chen et al.· 0 citations
Federated LLM fine-tuning enables large models to be adapted using private and geographically distributed data at the network edge, creating recurring and deadline-sensitive communication workloads across access and transport networks. This challenge is particularly relevant in mobile RANs, where wireless variability,...
E. Paolini, Andrea Pinto, Flavio Esposito et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 7, 2026
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduOct 6, 2026