Skip to content
Open access

A Constraint-Aware Proximal Policy Optimization Method for Task Scheduling and Dynamic Reconfiguration of Maritime Unmanned Systems

Oct 2026 · Journal of Marine Science and Engineering · 0 citations

Abstract

To address slow dynamic-reconfiguration response and difficulty in satisfying hard constraints in the task scheduling of heterogeneous maritime unmanned platforms, this paper proposes a constraint-aware proximal policy optimization (CPPO). CPPO models task scheduling as a Markov decision process; applies a three-layer action mask filtering invalid actions by platform survival and capacity, communication-link reachability, and payload capability matching to guarantee hard constraints; and introduces a staged penalty adjustment strategy to balance exploration and constraint internalization. In eight scenarios spanning 14 to 100 platforms under static and platform-failure settings, CPPO is benchmarked against Greedy, improved NSGA-II, and LSTM-PPO. Ablations show the mask is the dominant factor: it raises the task completion rate from about 25% to 53–97% and the constraint satisfaction rate from about 0.1 to 0.65–1.0, while the staged penalty adds gains. Compared with Greedy, CPPO shows a statistically equivalent completion rate (p > 0.05) and a significantly higher constraint satisfaction rate in small-scale scenarios (p = 0.006), while its inference latency remains about 1 ms, 2–3 orders of magnitude faster. Compared with LSTM-PPO, recurrent memory brings no gain in quality or constraint satisfaction, while training time increases 2–7-fold, validating the lightweight MLP encoder. CPPO provides an “offline training–online inference” solution for maritime unmanned system reconfiguration.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.