DC‐MAPPO is developed, which introduces a Directional Clamp mechanism to improve policy update stability under high‐concurrency conditions and a dynamic action masking strategy is designed to ensure decision feasibility and reduce invalid exploration.
Abstract
In edge–cloud collaborative computing, efficient scheduling of concurrent task chains is essential for reducing end‐to‐end latency. However, heterogeneous resources, dynamic link contention, and coupled routing‐computation decisions make it difficult to improve system efficiency while reducing communication conflicts. To address these challenges, we formulate concurrent heterogeneous task‐chain scheduling under partial observability as a multi‐agent partially observable Markov game and propose a MARL‐based scheduling method. Each task chain is represented by a mobile agent that makes routing, computation, data‐access, and waiting decisions based on local observations. Building upon a multi‐agent proximal policy optimization (MAPPO) architecture, we develop DC‐MAPPO, which introduces a Directional Clamp mechanism to improve policy update stability under high‐concurrency conditions. In addition, a dynamic action masking strategy is designed to ensure decision feasibility and reduce invalid exploration. Experimental results across multiple network topologies show that the proposed method outperforms the compared baselines in most tested settings. Under high‐load conditions, DC‐MAPPO reduces system makespan by 10.7%–17.1% and reduces the link reservation failure rate by 8.4–13.2 percentage points compared with MAPPO.
This paper addresses the joint task offloading and resource allocation problem in multi-user MEC systems and proposes a decentralized control framework based on Multi-Agent Reinforcement Learning (MARL), which achieves lower total system cost and faster convergence than the full-local, full-offload, and heuristic baselines.
Youssef Oukissou, Mohamed Amine Meddaoui, Ayoub Belaidi et al.· International journal of Com...· 0 citations
High-velocity workloads and intricate task dependencies inherent in distributed stream-processing systems pose a fundamental challenge to efficient resource allocation. Traditional heuristic and single-agent reinforcement learning (RL) schedulers frequently fail to recognize these complex network and data-flow interactions, leading to severe resource fragmentation and catastrophic tail latency spikes. In order to accomplish coordinated, low-latency scheduling, we propose a Topology-Aware Multi-Agent Reinforcement Learning (TAMARL) framework utilizing a Centralized Training and Decentralized Execution (CTDE) architecture. TAMARL allows distributed agents to optimize task placement across heterogeneous cluster nodes and prevent backpressure cascades by integrating topology-aware state representations. We evaluate TAMARL on a production-grade cloud-native stack leveraging Apache Flink and Kubernetes. Compared to state-of-the-art baselines across six demanding stress-test scenarios, experimental evaluations demonstrate that TAMARL improves Service Level Objective (SLO) attainment by 27% while reducing P99 tail latency by up to 68%. Additionally, TAMARL maintains stable, resilient performance under 90% cluster utilization while securing 95% network locality.
Sunday J. Awine, Jinwei Liu· IEEE International Conferenc...· 0 citations
Results provide initial evidence that multi-round CNP refinement is the principal protocol-level gain, with LLM assistance adding value for qualitative and uncertain runtime context.
This paper proposes a multi-agent reinforcement learning (MARL) framework for TSN scheduling, where each TSN queue is modeled as an autonomous agent and the Heterogeneous-Agent Proximal Policy Optimization (HAPPO) algorithm is employed to explicitly model inter-agent dependencies and jointly optimize service delivery across queues.
Marcos Carvalho, Fatih Temiz, Shavbo Salehi et al.· 0 citations
A constraint-aware multi-agent edge collaborative offloading algorithm (CARE-CTDE) that achieves better scheduling performance, resource utilization, and constraint satisfaction than baseline methods in dynamic heterogeneous MEC scenarios, demonstrating its effectiveness and robustness for constrained edge computing systems.
Yuxuan Yang, Hexing Wang, Yang Zhou· Mathematics· 0 citations
Edge-assisted cognitive radio networks require efficient scheduling mechanisms to jointly manage opportunistic spectrum access, task offloading, energy consumption, and latency constraints. Existing multi-agent scheduling approaches often rely on fixed penalty terms or average queue-based constraints, which may not effectively control service-level violations under uncertain spectrum availability and dynamic edge-resource contention. This work proposes SCOPE, a Safe Causal-Graph Primal–Dual multi-agent scheduling framework for energy- and latency-constrained edge-assisted cognitive radio networks. The major strength of SCOPE is its integrated design, where belief-state augmentation improves decision-making under imperfect spectrum sensing, dual-relational causal graph coordination separately models’ interference coupling and computation-resource contention, and CVaR-based primal–dual optimization regulates tail-risk violations of latency and energy constraints. The framework follows a centralized-training and decentralized-execution structure, enabling coordinated learning during training while supporting scalable decentralized scheduling during deployment. Simulation results under dynamic user mobility, stochastic task arrivals, and varying primary-user activity show that SCOPE improves latency, energy efficiency, service-level constraint satisfaction, and throughput compared with existing scheduling methods. Ablation analysis further confirms the individual contribution of belief modeling, graph coordination, and risk-sensitive constraint enforcement.
T. Kannan, M. Lavanya, A. Ponraj et al.· Scientific Reports· 0 citations