ML-Based Hierarchical Prediction for Practical Energy Scheduling in Dynamic NTN-WPT Systems
With advancements in long-distance wireless power transfer (WPT) and space-based energy technologies, the integration of WPT into non-terrestrial networks (NTNs), hereafter referred to as NTN-WPT, is emerging as a promising approach for next-generation wireless networks. This paper proposes an energy-scheduling approach to jointly optimize energy efficiency, task completion rate, and task waiting time for power transfer from low Earth orbit satellites to terrestrial mobile user devices (UDs). To address the significant energy-scheduling challenges arising from satellite and UD mobility and further exacerbated by channel uncertainty due to stochastic propagation effects, we decompose the problem into three subproblems corresponding to a three-layer predictive framework: 1) a state prediction layer forecasts UD and satellite states; 2) an interaction mapping layer, employing a graph neural network (GNN), models the energy transfer efficiency between them; and 3) a decision-making layer determines the optimal energy allocation plan. We employ distinct machine learning (ML) methods within this framework, tailored to the specific requirements of each layer. Furthermore, balancing these competing objectives presents a challenging multi-objective optimization problem (MOP). We address this by adopting a key multi-objective reinforcement learning (MORL) technique: scalarizing the objectives into a single weighted-sum reward function. This scalarization transforms the MOP into a tractable, single-objective problem for the agents to solve. To help the agents balance these competing objectives effectively, we introduce a multi-agent deep learning model that integrates a self-attention mechanism with multi-agent proximal policy optimization (MAPPO). This approach provides a robust and efficient solution for WPT in NTNs, particularly for mission-critical scenarios. Simulation results show that the proposed approach can achieve a better overall trade-off than the baseline methods, maintaining competitive task completion rates and energy efficiency while reducing task waiting times. It also demonstrates robust performance under highly variable conditions.