Transformer-Based Offline Reinforcement Learning for Intelligent Torpedo Evasion in Autonomous Underwater Vehicles
Abstract
This paper proposes an offline reinforcement learning framework that combines a transformer architecture with conservative Q-learning (CQL) for torpedo evasion of autonomous underwater vehicles (AUVs). Conventional approaches that feed only a single-step observation into a multilayer perceptron (MLP) struggle to capture the temporal context that governs underwater engagements, such as the pursuit history of an incoming torpedo and the gradual deflection induced by acoustic decoys. The proposed method exploits the self-attention mechanism of a transformer encoder to precisely infer the approach patterns of torpedoes from sequences of past observations. Moreover, the policy is optimized exclusively on a pre-collected static dataset, which removes the cost and the physical risk of trial-and-error interactions, while the conservative penalty of CQL reduces value overestimation and discourages unsupported actions, which improves the empirical safety of the learned policy without constituting a formal safety guarantee. The central contribution lies not in the individual building blocks but in their unified design, in which sequential-context encoding, conservative offline optimization, and the launch timing of self-propelled acoustic decoys are cast into a single continuous action space dedicated to torpedo evasion, a combination that has not been jointly addressed in prior work. Experiments in a three-dimensional maritime simulator built on a six-degree-of-freedom vehicle model demonstrate that the proposed algorithm achieves an 88.2% mission success rate, outperforming a conventional MLP-based CQL baseline by 30.2 percentage points, and that it also exceeds the strongest modern offline RL baseline, implicit Q-learning (IQL), by 5.2 percentage points. In addition, multi-head attention heatmap analysis reveals that the agent autonomously learns a role-division mechanism over spatiotemporal contexts, and trajectory analysis confirms that the agent acquires a composite tactic of rapidly increasing its depth and deploying self-propelled acoustic decoys immediately after torpedo detection. These results indicate that the proposed framework offers a conservative and computationally deployable offline learning approach for safety-critical underwater autonomous systems, while the present validation is confined to a custom simulation environment, so that operational deployment remains contingent on the hardware-in-the-loop testing, field trials, and a dedicated sim-to-real study that are left as future work.