Skip to content

An LLM-Driven Hybrid Online-Offline MARL Architecture for AUV Cluster Network to Enable Smart Underwater Target Encirclement

Oct 2026 · IEEE wireless communications · Vol 33, pp. 137-144 · 2 citations · 18 references

Abstract

With the rapid advancement of underwater networking and multi-agent reinforcement learning (MARL) technologies, autonomous underwater vehicle (AUV) cluster networks have emerged as a promising framework for enabling smart underwater missions, particularly in cooperative target encirclement. However, training MARL policies in such environments faces a critical challenge: acquiring large-scale interaction data is infeasible due to the severely constrained communication bandwidth, especially for online MARL frameworks. Therefore, achieving fast policy convergence with limited data collection is essential for practical deployment. This article proposes an LLM-Empowered Hybrid Training (LLM-EHT) architecture for MARL, which leverages the reasoning and generative capabilities of large language models (LLMs) to facilitate the transition from online to offline MARL under limited-sample conditions. Specifically, the proposed architecture first collects a small amount of online interaction data. It then employs an LLM to synthesize an offline dataset, guided by a dedicated alignment loss that enforces consistency between policy actions and LLM-generated references. This process significantly accelerates MARL policy convergence. Building upon LLM-EHT, we further develop a smart and cooperative underwater target encirclement scheme, incorporating a scalable state space representation method and a smart target encirclement policy to ensure robust, efficient, and scalable operation. Evaluation results showcase that the proposed scheme significantly reduces the required sample size while achieving faster convergence and maintaining stable encirclement effectiveness.

View source

Similar papers

2026

Hierarchical Multi-Agent Reinforcement Learning for Networked Multi-AUV Data Collection in UWSNs

Underwater wireless sensor networks form the foundation of the marine Internet of Things, but timely data delivery remains challenging in deep and remote deployments. Although AUV-assisted data collection reduces reliance on energy-constrained multi-hop acoustic relays, slow vehicle mobility and repeated surfacing for...

Yu-Hui Han, Peng Jin, Xiao Huang et al. · 0 citations
Open access Aug 2026

SkyAgent: A lightweight LLM-driven reinforcement learning framework for adaptive cooperative path planning of two UAVs

This work provides a feasible technical pathway and reproducible evaluation benchmark for the collaborative deployment of lightweight LLM planner, sub-goal guidance, sensor observations, cooperative reward, and reward shaping components and quantifies the indispensability of the LLM planner.

Yuting Cao, Zheng Zhao, Jiekai Wu et al. · 0 citations

Physics-prior-driven distributed deep reinforcement learning for multi-UAV path planning

A physics-prior-driven decentralized deep reinforcement learning (DRL) framework Functioning as a scalable distributed computing paradigm via decentralized training with decentralized execution (DTDE), the framework mitigates the curse of dimensionality.

Jian-Hua Liu, Hai-Tao Zhou, Jia-Jia Liu et al. · 0 citations
Conference Sep 2026

Toward Agentic Intelligence in Non-Terrestrial Networks: A Roadmap for Federated and Quantum-Enhanced Learning

Non-terrestrial networks (NTNs), comprising satellites, uncrewed aerial vehicles (UAVs), and high-altitude platform stations (HAPS), are key enablers of sixth-generation (6G) wireless systems, providing global seamless connectivity. However, NTN environments present challenges, including high propagation delays, interm...

Tajwar Sattar, Najam Ul Hasan, Waleed Ejaz · 0 citations
Preprint Aug 2026

Multi-AUV Ad-hoc network-based Target Tracking: A Value Gradient Guidance Multi-Agent Diffusion Reinforcement Learning Approach

Experimental results show that VGG-MADiffRL consistently achieves faster convergence, higher tracking accuracy, and smoother training dynamics in cooperative tracking scenarios, validating its effectiveness and practical engineering value in dynamic underwater settings.

Jiaao Ma, Chuan Lin, Guang-Jie Han et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.