Oct 2026· IEEE wireless communications· Vol 33, pp. 137-144· 2 citations· 18 references
Abstract
With the rapid advancement of underwater networking and multi-agent reinforcement learning (MARL) technologies, autonomous underwater vehicle (AUV) cluster networks have emerged as a promising framework for enabling smart underwater missions, particularly in cooperative target encirclement. However, training MARL policies in such environments faces a critical challenge: acquiring large-scale interaction data is infeasible due to the severely constrained communication bandwidth, especially for online MARL frameworks. Therefore, achieving fast policy convergence with limited data collection is essential for practical deployment. This article proposes an LLM-Empowered Hybrid Training (LLM-EHT) architecture for MARL, which leverages the reasoning and generative capabilities of large language models (LLMs) to facilitate the transition from online to offline MARL under limited-sample conditions. Specifically, the proposed architecture first collects a small amount of online interaction data. It then employs an LLM to synthesize an offline dataset, guided by a dedicated alignment loss that enforces consistency between policy actions and LLM-generated references. This process significantly accelerates MARL policy convergence. Building upon LLM-EHT, we further develop a smart and cooperative underwater target encirclement scheme, incorporating a scalable state space representation method and a smart target encirclement policy to ensure robust, efficient, and scalable operation. Evaluation results showcase that the proposed scheme significantly reduces the required sample size while achieving faster convergence and maintaining stable encirclement effectiveness.
Underwater wireless sensor networks form the foundation of the marine Internet of Things, but timely data delivery remains challenging in deep and remote deployments. Although AUV-assisted data collection reduces reliance on energy-constrained multi-hop acoustic relays, slow vehicle mobility and repeated surfacing for...
Yu-Hui Han, Peng Jin, Xiao Huang et al.· IEEE Transactions on Cogniti...· 0 citations
This work provides a feasible technical pathway and reproducible evaluation benchmark for the collaborative deployment of lightweight LLM planner, sub-goal guidance, sensor observations, cooperative reward, and reward shaping components and quantifies the indispensability of the LLM planner.
Yuting Cao, Zheng Zhao, Jiekai Wu et al.· Journal of King Saud Univers...· 0 citations
A physics-prior-driven decentralized deep reinforcement learning (DRL) framework Functioning as a scalable distributed computing paradigm via decentralized training with decentralized execution (DTDE), the framework mitigates the curse of dimensionality.
Jian-Hua Liu, Hai-Tao Zhou, Jia-Jia Liu et al.· Journal of Supercomputing· 0 citations
Experimental results show that VGG-MADiffRL consistently achieves faster convergence, higher tracking accuracy, and smoother training dynamics in cooperative tracking scenarios, validating its effectiveness and practical engineering value in dynamic underwater settings.
Jiaao Ma, Chuan Lin, Guang-Jie Han et al.· 0 citations