Skip to content
Open access

An Adaptive Scheduling Algorithm Integrating Hierarchical Reinforcement Learning and Semi-Markov Decision Processes

Jul 2026 · Applied Sciences · Vol 16, pp. 6570 · 0 citations · 20 references

TL;DR

Simulation results in multi-constraint environments demonstrate that the Adaptive Hierarchical Semi-Markov Decision Process (AH-SMDP) framework effectively improves scheduling performance compared to standard MAPPO and PPO algorithms.

Abstract

Coordinating multiple unmanned aerial vehicle (UAV) systems under strict energy and temporal constraints remains a complex scheduling problem. Existing reinforcement learning methods typically rely on fixed-time-step modeling, which struggles to accommodate flight actions of varying durations and often leads to temporal mismatches between task planning and physical execution. To address this limitation, we propose an Adaptive Hierarchical Semi-Markov Decision Process (AH-SMDP) framework. This architecture decouples task allocation from execution by modeling variable-length actions via an SMDP. An event-driven synchronization mechanism is introduced to align the swarm’s decision-making rhythm with actual task completion times. Additionally, a state-aware reward formulation and a dynamic action space pruning strategy are designed to help UAVs balance energy efficiency with deadline compliance. Simulation results in multi-constraint environments demonstrate that the AH-SMDP framework effectively improves scheduling performance compared to standard MAPPO and PPO algorithms. Under the evaluated experimental settings, the proposed method yields improvements of approximately 30% in average task completion rate, 40% in energy reduction, and 60% in convergence stability. Ablation studies further suggest that this integrated framework offers a viable and effective approach for multi-UAV scheduling.

Read PDF

Similar papers

2026

Safe Reinforcement Learning for Autonomous and Constraint-Aware Earth Observation Satellite Scheduling

Abstract. Satellite scheduling requires balancing dynamic tasking demands, limited resources, and real-time adaptability, conditions under which traditional rule-based and heuristic methods often fall short. This study explores Safe Reinforcement Learning (SRL) as a framework for autonomous, adaptive, and constraint-aware mission planning. We propose a frugal fine-tuning strategy that adapts a Proximal Policy Optimization (PPO) baseline to new operational domains within the BSK-RL simulation environment, leveraging soft constraints modeled through cost signals. Two SRL approaches are evaluated: (i) reward reshaping, which integrates constraint costs into the reward function, yielding stability improvements and an 8% reduction in actions per episode without sacrificing cumulative rewards; and (ii) Lagrangian-based Constrained Policy Optimization (CPO), which separately optimizes rewards and constraint costs, prioritizing constraint satisfaction and reducing actions by 20–28% at the expense of overall reward. Results demonstrate that SRL enhances both efficiency and constraint adherence in scheduling, while the proposed fine-tuning approach enables rapid transfer to new mission scenarios. These findings highlight SRL as a promising pathway toward scalable, safe, and adaptive satellite operations.

Fabio Favaretto · 0 citations
Preprint Aug 2026

LLM-Based Hierarchical Coordinated Control with Continuation-Aware Policy Learning

Coordinating multiple interacting units in complex engineering systems is challenging when system interactions are difficult to model, operational information is heterogeneous, and low-level actions must satisfy strict constraints. We propose an LLM-based hierarchical framework in which the LLM coordinates interacting units based on heterogeneous operational context, while task-specific controllers or optimizers generate executable and constraint-aware actions. We further introduce Continuation-Aware GRPO to capture the consequences of coordination decisions over subsequent control intervals. Rather than judging a decision only by its immediate outcome, the method also evaluates how the system evolves afterward under the current policy. We validate the framework on multi-ramp traffic control and virtual power plant (VPP) energy management, using simplified system models for training and more realistic simulators for evaluation. Across both tasks, the proposed method consistently outperforms direct task-specific control and optimization, end-to-end reinforcement learning, rule-based and RL-based hierarchical coordination, and prompting-only LLM coordinators, demonstrating the value of heterogeneous-context reasoning, hierarchical execution, and continuation-aware policy learning.

Changhong He, Jinda Gao, Xinkuan Liu et al. · 0 citations
Open access Jul 2026

An Experience-Guided MAPPO Framework for Multi-UAV Cooperative Tracking in Continuous Action Spaces

A cooperative guidance law based on the experience-guided multi-agent proximal policy optimization (E-MAPPO) algorithm is proposed for multiple unmanned aerial vehicles (UAVs) to track dynamic points of interest in civilian applications, such as collaborative search and rescue and environmental monitoring. In multi-UAV cooperative tracking, accurate arrival-time coordination is important for improving collaborative task execution, but it remains challenging because of continuous action spaces, target maneuvering, uncertain time-to-go estimation, and inefficient exploration in multi-agent reinforcement learning. Specifically, a multi-UAV cooperative guidance environment is formulated, and the problem is modeled as a Markov decision process. To address the challenges of large action spaces and poor convergence in multi-agent reinforcement learning, an experience-guided MAPPO framework is introduced to enhance training efficiency and policy stability. Different from standard MAPPO, the proposed E-MAPPO introduces proportional-navigation-guided experience only during the early training stage to guide exploration, while the final policy is still optimized through the MAPPO objective. Subsequently, a composite reward function is designed by integrating distance-based heuristic terms with auxiliary guidance signals, thereby improving exploration efficiency and facilitating coordinated rendezvous and tracking of dynamic references. Comparative simulations with cooperative proportional navigation guidance (CPNG), sliding mode control (SMC), and standard MAPPO are conducted under different target motion scenarios. The results show that E-MAPPO reduces the average convergence step by 17.07% compared with MAPPO. In the straight-moving target scenario, E-MAPPO reduces the cooperative time error by 55.10% compared with CPNG and by 8.33% compared with MAPPO. In the S-type maneuvering target scenario, E-MAPPO reduces the cooperative time error by 55.81% compared with CPNG and by 9.52% compared with MAPPO. Monte Carlo experiments further verify its effectiveness and robustness. Additional robustness tests under Gaussian measurement noise, observation bias, and communication delay show that the proposed method maintains acceptable tracking accuracy and cooperative timing performance under different uncertainty conditions. In addition, the results indicate that the proposed method generalizes well to different types of maneuvering targets.

Hao Xiong, Minghu Tan, Xiaoyu Liu et al. · 0 citations
2026

Energy-Aware Multi-UAV Collaboration for Data Collection and Trajectory Planning With MADDPG

Unmanned Aerial Vehicles (UAVs) are pivotal for facilitating data collection in emergency scenarios. Despite the potential of Multi-Agent Deep Reinforcement Learning (MADRL) in coordinating such systems, existing researches struggle to resolve the high-dimensional coupling of data collection, trajectory planning, and energy scheduling under strict collision avoidance and Return-To-Base (RTB) constraints. This paper proposes a energy-aware cooperative MADRL framework designed to maximize data collection utility under energy constraints. Specifically, we employ a Multi-Agent Deep Deterministic Policy Gradient (MADDPG) approach featuring a Centralized Training with Decentralized Execution (CTDE) design and a multi-objective reward mechanism to balance conflicting optimization goals. Extensive simulations validate the advantages of the proposed framework over leading baselines. Notably, the algorithm exhibits significant quantitative advantages in complex high-load scenarios. These outcomes prove that our method achieves higher task completion rates while strictly adhering to RTB and safety protocols.

Jing Mei, Jinglei Xu, Zhao Tong et al. · 0 citations
Open access Jul 2026

A Multi-UAV Cooperative Mission Planning Method Based on Multi-Agent Guided Soft Actor–Critic

A multi-agent guided soft actor–critic (MAGSAC) deep reinforcement learning algorithm to enable multiple UAVs to simultaneously arrive at multiple constant-velocity moving targets and outperforms existing mainstream algorithms in synchronization success rate, temporal synchronization accuracy, and safety.

Shuanli Jia, Naiming Qi, Zheng Li et al. · 0 citations
#reinforcement learning Open access Aug 2026

Multi-Objective Task Allocation and Path Planning in Heterogeneous Multi-Robot Systems Using Hierarchical Framework and Reinforcement Learning

This study proposes a multi-objective optimization framework for task allocation and path planning in transport-oriented multi-robot systems. The framework explicitly considers heterogeneous robot capabilities and load capacities while jointly minimizing task completion time and overall energy consumption. A hierarchical architecture is adopted, consisting of two stages. In the upper layer, the NSGA-II algorithm evaluates task allocation strategies and constructs a Pareto-optimal solution space, enabling decision-makers to select solutions according to optimization preferences or operational constraints. In the lower layer, deep neural networks and reinforcement learning are employed for multi-agent learning to generate collision-free paths for the assigned tasks. This hierarchical design enables capability-aware task allocation while providing flexibility to accommodate optimization priorities. Simulation and experimental results demonstrate that the proposed framework effectively addresses complex scenarios involving task dependencies, improves path-learning efficiency and task allocation performance, and provides multiple interpretable trade-off solutions without compromising single-objective performance. These results highlight the framework’s effectiveness, scalability, and practical applicability to real-world multi-robot transportation tasks.

Sheng-Hsiang Luo, Rongshun Chen · 0 citations