Skip to content

Safe Reinforcement Learning for Autonomous and Constraint-Aware Earth Observation Satellite Scheduling

2026 · Materials Research Proceedings · 0 citations

Abstract

Abstract. Satellite scheduling requires balancing dynamic tasking demands, limited resources, and real-time adaptability, conditions under which traditional rule-based and heuristic methods often fall short. This study explores Safe Reinforcement Learning (SRL) as a framework for autonomous, adaptive, and constraint-aware mission planning. We propose a frugal fine-tuning strategy that adapts a Proximal Policy Optimization (PPO) baseline to new operational domains within the BSK-RL simulation environment, leveraging soft constraints modeled through cost signals. Two SRL approaches are evaluated: (i) reward reshaping, which integrates constraint costs into the reward function, yielding stability improvements and an 8% reduction in actions per episode without sacrificing cumulative rewards; and (ii) Lagrangian-based Constrained Policy Optimization (CPO), which separately optimizes rewards and constraint costs, prioritizing constraint satisfaction and reducing actions by 20–28% at the expense of overall reward. Results demonstrate that SRL enhances both efficiency and constraint adherence in scheduling, while the proposed fine-tuning approach enables rapid transfer to new mission scenarios. These findings highlight SRL as a promising pathway toward scalable, safe, and adaptive satellite operations.

View source