Skip to content

Scaling SDN Control Planes with Multi-Agent Reinforcement Learning

Sep 2026 · International Symposium on Networks, Computers and Communications · pp. 1-5 · 0 citations · 13 references

Abstract

Modern Wide-Area Software-Defined Networks (WAN-SDNs) face critical challenges regarding scalability, control-plane propagation latency, and single-points-of-failure due to their reliance on centralized control architectures. While logically distributed SDN controllers mitigate structural vulnerabilities, existing static or heuristic-based controller loadbalancing and inter-domain routing protocols cannot adapt to highly dynamic, asymmetric traffic demands. This paper presents an Intelligent, Distributed Control Plane framework driven by Multi-Agent Reinforcement Learning (MARL). In the proposed architecture, each distributed SDN controller operates as an autonomous deep Q-learning agent that continuously monitors local domain topologies and dynamically optimizes flow-rule allocations with neighboring controllers. To mitigate inter-controller communication overhead and control-channel flooding, we implement an Independent Double Deep Q-Network (IDDQN) variant augmented with a spatially discounted global reward function, explicitly preventing localized routing loops and policy divergence. We emulate our architecture on the Mininet platform utilizing the standard Internet2 topology. Empirical results demonstrate that our MARL-driven framework reduces average flow-setup latency by up to 28%, enhances network throughput by 19%, and accelerates fault-recovery times by 41% compared to legacy centralized control models, static distributed routing benchmarks, and standard MARL baselines, presenting a resilient, self-optimizing control topology for next-generation distributed networks.

View source

Similar papers

Conference Aug 2026

Agentic AI for QoS-driven Adaptive Routing in Small-to-Medium Scale Software-Defined Networks

Software-Defined Networking enables centralized, programmable traffic management, yet existing routing approaches face limitations in balancing multiple quality of service objectives. Classical algorithms minimize single metrics, while Machine Learning and Deep Learning methods require extensive training data and lack...

Robin E. Valenzuela, Alonica R. Villanueva · 0 citations
Open access Aug 2026

Design of a Scalable Control Plane for Large-Scale SDNs

Findings affirm that the suggested scalable control plane is practical in supporting large scale SDN implementation and is therefore applicable in future carrier grade, data center and wide area network deployments at realistic workloads with varying topological setups in the modern programmable networks in the world.

A. Nagadeepan, Vishakha Abhay Gaidhani, Bhambare Rajesh et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Prior-Guided Heuristics to Deployable Agents: Accelerating Demonstration-Driven Reinforcement Learning for Deadline-Constrained Network Control

This paper introduces Effective Congestion (EC), a deadline-aware metric family that quantifies interface congestion by packet urgency and proactively filters non-viable traffic, coupled with a Uniform Path Grouping (UPG) distribution heuristic promoting robust load-balancing; the resulting policies are embedded into M...

Vincenzo Norman Vitale, Mohammad Solki, A. Tulino et al. · 0 citations
Open access Sep 2026

Federated learning for edge devices delay control in software defined wide area networks

An SDN-orchestrated architecture for delay control in wide-area networks, combining edge-based GRU Active Queue Management (AQM) with layer-wise federated averaging, indicates that federated averaging reduces rare congestion events without meaningful degradation of normal bottleneck operation.

Karol Marszałek, Adam Domański · 0 citations
Preprint Sep 2026

AoI-Driven Hierarchical Learning for Cooperative Resource Sharing in Multi-Operator UAV Networks

Uncrewed aerial vehicle (UAV)-assisted networks provide a versatile paradigm for on-demand connectivity. However, in multi-operator aerial networks (MOANs), the joint optimization of cooperative resource sharing and 3D trajectory control to maintain information freshness is a complex combinatorial problem, which can be...

Atefeh Hajijamali Arani, M. Shirvanimoghaddam, A. Mehbodniya et al. · 0 citations
Open access 2026

Lyapunov-DLD-Based Latency and Power Optimization in 5G O-RAN for Federated Learning

Experimental results demonstrate that the proposed framework improves convergence, accuracy, scalability, and signal-to-noise ratio, data rate, while simultaneously reducing latency, and energy consumption for both CIFAR-10 and FEMNIST datasets compared with the FedProx, FedADMM, LyFeD and FL-MEC benchmark schemes.

Kofi Kwarteng Abrokwa, Qi Jiang, Zhou-Qin Ma et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.