The proposed multi-agent reinforcement learning policy attains slightly higher throughput with fewer handovers by offloading a fraction of the users to the MEO and GEO layers, an emergent multi-orbit behavior that drives its favorable throughput and handover trade-off.
Abstract
Future sixth-generation non-terrestrial networks are expected to combine low Earth orbit (LEO), medium Earth orbit (MEO), and geostationary Earth orbit (GEO) satellites, whose complementary layers must be coordinated through joint user association, power allocation, and handover management under fast LEO dynamics. This paper studies this problem by formulating it as a mixed-integer nonlinear program and decomposing it into a multi-agent reinforcement learning (MARL) policy that selects the associations and a convex power-allocation subproblem solved exactly at each time slot that defines the reward of the MARL part. The association policy is trained with multi-agent proximal policy optimization (MAPPO) and the targeted multi-agent communication (TarMAC) mechanism, and is made aware of the orbital layer through a state that encodes layer-dependent handover penalties. Evaluated on a realistic multi-constellation scenario built from real two-line element data over Nairobi, Kenya, the proposed policy reaches 92% of the throughput of a greedy signal-to-noise ratio (SNR) maximizing scheme while triggering more than four times fewer handovers, and improves throughput by roughly 14% over a conservative stay heuristic. Compared to an LEO-only learned policy of identical architecture, it attains slightly higher throughput with fewer handovers by offloading a fraction of the users to the MEO and GEO layers, an emergent multi-orbit behavior that drives its favorable throughput and handover trade-off.
Low Earth Orbit (LEO) constellations are required to process increasing volumes of heterogeneous tasks from ground networks. Intermittent inter-satellite links, heterogeneous onboard resources, and time-varying traffic loads make collaborative task offloading difficult for static or reactive strategies. Although multi-agent reinforcement learning (MARL) provides an adaptive solution, existing model-free MARL methods often suffer from slow convergence, insufficient foresight, and limited robustness in dynamic satellite environments. To address these challenges, this paper proposes a World Model-Based Multi-Agent Proximal Policy Optimization (WM-MAPPO) framework for space computing power networks. The offloading problem is formulated as a partially observable multi-agent decision-making process, where LEO satellites make decentralized decisions under incomplete local observations. A predictive world model learns latent transition dynamics of network states and provides future context for proactive planning. Meanwhile, a Transformer-based policy architecture captures inter-agent dependencies and supports cooperative scheduling under centralized training and decentralized execution (CTDE). Simulation results show that WM-MAPPO achieves higher task completion ratios, lower average latency, improved energy efficiency, and stronger robustness than model-free MARL baselines, heuristic methods, Lyapunov-based scheduling, and MINLP-inspired optimization.
Yuqi Cong, Zhiwei Wei, Jiarui Chen et al.· IEEE Transactions on Cogniti...· 0 citations
Low Earth orbit (LEO) satellite-terrestrial communication systems grapple with significant challenges posed by their inherent dynamism and substantial transmission delays. To address these critical issues, this paper proposes a novel hybrid-medium transmission optimization framework that leverages high-altitude platforms (HAPs) as relays. Our primary objective is to minimize end-to-end system delay through the joint optimization of transmission mode selection and wireless communication resource allocation. The resulting joint optimization problem is formulated as a computationally intractable mixed-integer nonlinear programming (MINLP). We present a hierarchical solution strategy to tackle this complexity. Firstly, Lagrangian optimization is employed to analytically derive the intrinsic coupling between resource allocation and transmission mode selection, thereby simplifying the problem into a sequential decision-making process. This sequential problem is subsequently framed as a Markov decision process (MDP), enabling the design of a deep reinforcement learning (DRL) agent tasked with dynamically learning the optimal transmission mode selection policy. By maximizing cumulative long-term rewards, our DRL-based approach effectively reduces overall system delay, unlocking enhanced performance potential for future 6G networks.
Yi Huang, Jin Li, Yanwen Zhu et al.· IEEE Transactions on Communi...· 0 citations
In Advanced Air Mobility (AAM) applications, jointly optimizing multi-agent motion control along predefined flight routes and spectrum access is highly challenging due to the tight coupling among mobility, interference, and safety constraints under limited spectrum resources. This paper proposes a Large Language Model (LLM)-guided cooperative decision-making framework for joint velocity control and bidirectional channel selection in an AAM system with Aerial Vehicles (AVs) communicating with ground Base Stations (BSs) while following predefined linear routes. We formulate the problem as a cooperative Markov game with a discrete action space that includes both velocity selection and uplink and downlink channel access, while satisfying Signal-to-Interference-plus-Noise Ratio (SINR) quality requirements and collision avoidance constraints. To obtain reliable expert behavior, we first learn a near-optimal policy using Multi-Agent Reinforcement Learning (MARL) with Value Decomposition Dueling Double Deep Q-Networks (VD3QN). We then treat joint decision-making as a sequence generation task and employ Large Language Models (LLMs) to generate complete joint action sequences from structured environment descriptions, under both Prompt Engineering (PE) and Parameter-Efficient Fine-Tuning (PEFT) via Low-Rank Adaptation (LoRA) on expert demonstrations. Extensive simulations show that structured prompting improves decision quality, while LoRA fine-tuning further increases reward, reduces variance, and yields decisions that closely match the expert policy. Beyond this imitation role, the LLM layer turns 6-AV expert demonstrations into a sequence-level decision generator. In an unseen 10-AV scenario, this generator achieves stronger zero-shot generalization than the VD3QN policy transferred from the 6-AV environment.
Qingyang Li, Adnan Quadri, Hongxiang Li et al.· IEEE Access· 0 citations
Maritime satellite communications (SATCOMs) are expected to support high-capacity ship-to-satellite uplinks for remote maritime services beyond terrestrial coverage, with low-Earth-orbit (LEO) satellites providing wide-area connectivity. However, robust uplink beamforming in LEO maritime SATCOMs is challenging because dynamic ship–satellite geometry, wave-induced attitude motion, imperfect channel state information, and multiship interference make transmit power, ship-side transmit beamforming, and satellite-side receive combining tightly coupled. Accordingly, we formulate a long-term spectral efficiency (SE) maximization problem under transmit-power and quality-of-service constraints. An attitude-aware uplink channel model is developed by incorporating roll, pitch, and yaw motions into the effective angle-of-departure/angle-of-arrival evolution. Based on this model, the problem is cast as a heterogeneous decentralized partially observable Markov decision process. We then propose a robust heterogeneous cooperative QMIX (RHC-QMIX) framework under centralized training and decentralized execution, where type-specific recurrent local Q-networks, history-refined angular features, and centralized monotonic value mixing coordinate ship and satellite agents. Extensive simulations demonstrate that in the load-controlled scalability evaluation, RHC-QMIX achieves an average network SE of 25.32 bps/Hz, improves over alternating optimization by up to 51.00% as the satellite load increases, and outperforms heterogeneous cooperative QMIX by 16.38% on average under network-size scaling; it also maintains more stable SE under severe sea-state-induced ship motion.
Low Earth orbit (LEO) satellite communication, able to provide ubiquitous and continuous connectivity, has become a vital component of future sixth-generation global networks. To further improve service continuity and support spatially non-uniform traffic demand, LEO satellite systems are evolving toward dense multi-shell constellations, where overlapping coverage enables multiple satellites to serve the same traffic region. However, under limited onboard beams, spectrum, and power budgets, such overlap may lead multiple beams to request the same traffic cell, causing duplicate beam-cell requests (DBRs) in downlink scheduling. To address this challenge, we propose a structured multi-agent reinforcement learning (MARL) framework based on multi-agent proximal policy optimization (MAPPO), termed ShellMean-MAPPO, for downlink resource allocation with explicit conflict resolution. Specifically, we first formulate a long-term scheduling problem that separates pre-resolution beam-cell requests from post-resolution retained transmissions, enabling DBRs to be modeled together with resource block (RB) chunk and power allocation. By leveraging compact local information and per-shell summaries, a typed encoder is then designed to capture heterogeneous service and contention states without relying on a full global map. Furthermore, an autoregressive policy is developed to generate beam-cell, RB chunk, and power-share decisions in accordance with the downlink scheduling sequence, while a deterministic replay-based post-resolution credit mechanism transforms team outcomes into per-beam training signals. Extensive simulation results validate the effectiveness of ShellMean-MAPPO, and demonstrate its advantages over representative MARL schemes in terms of scheduling performance and conflict mitigation.
Li Zhen, Qi-Hao Zhang, Qing-Zhi Meng et al.· IEEE Open Journal of the Com...· 0 citations
Low Earth Orbit (LEO) communication networks are an important component of non-terrestrial networks (NTNs) in sixth-generation (6G) communication systems. LEO satellites are characterized by low propagation delay and highly time-varying topology. Space-based mobility management can effectively reduce transmission delay; however, rapid network variation makes space-based control node deployment and reconfiguration difficult to model and solve. Focusing on dual-layer LEO Walker constellations, this paper investigates the joint optimization of control node deployment and dynamic reconfiguration, and formulates a 0–1 mixed-integer linear programming model with multiple practical constraints, aiming to minimize the total handover and migration delay. The model incorporates practical constraints such as the CN resource budget, unique management of access layer satellites, inter-layer reachability, feeder link connectivity, non-empty control node (CN) management, and onboard resource capacity. To support online deployment, we propose a rolling-horizon migration-aware dynamic greedy control node placement algorithm (RH-MA-DGCNP), which updates the CN placement and the affiliation between access-layer satellites and CNs at each reconfiguration epoch while jointly considering the handover delay and the migration delay caused by transferring control-affiliation states from previous serving CNs to new serving CNs. A comparison with exact current-epoch MILP solutions obtained by CPLEX on validation instances shows that RH-MA-DGCNP achieves small optimality gaps with shorter computation time. Simulation results show that RH-MA-DGCNP achieves the lowest mean handover delay among all benchmark schemes and the lowest cumulative total delay cost among the quasi-dynamic and dynamic benchmark schemes. The CDF of handover delay further indicates that RH-MA-DGCNP has a higher proportion of low-delay handover events and effectively suppresses extremely high-delay handover cases. Sensitivity analyses under different elevation angle thresholds and ground station deployments further show that RH-MA-DGCNP maintains its performance advantage over the benchmark schemes under different network settings.
Yang Liu, Wen Liu, Wenliang Lin et al.· Electronics· 0 citations