Sample-Efficient Autonomous Surface Vehicle Navigation with Enhanced Experience Replay
Abstract
This paper proposes a sample-efficient reinforcement learning framework for navigation of Autonomous Surface Vehicles (ASVs) in complex, obstacle-dense environments. Learning reliable navigation policies in such settings is challenging due to unsafe early exploration and high sample complexity. To address these issues, we integrate Distributional Reinforcement Learning with Deep Q-learning from Demonstrations (DQfD) and Count-Based Experience Replay (CbER), enabling effective learning from both expert priors and prioritized recent experiences. Furthermore, we critically evaluate the integration of Hindsight Experience Replay (HER) in navigational setups that already employ dense shaping rewards. Extensive simulations demonstrate that while HER offers limited marginal utility in the presence of strong shaping signals, the combined framework accelerates convergence by up to 6.3× compared to baselines.