Transformer-Based Offline Reinforcement Learning for Intelligent Torpedo Evasion in Autonomous Underwater Vehicles
This paper proposes an offline reinforcement learning framework that combines a transformer architecture with conservative Q-learning (CQL) for torpedo evasion of autonomous underwater vehicles (AUVs). Conventional approaches that feed only a single-step observation into a multilayer perceptron (MLP) struggle to captur...