Skip to content
Open access

PV-STAM: Velocity-Aware Attention for Mapless Deep Reinforcement Learning Navigation in Dynamic Environments

Sep 2026 · Applied Sciences · 0 citations · 20 references

Abstract

Mapless deep reinforcement learning (DRL) navigation in dynamic indoor environments is difficult under single-frame 2D LiDAR, which reports where an obstacle is but not whether it is approaching. We introduce the Positional-Velocity Spatio-Temporal Attention Module (PV-STAM), a compact perception block (19,968 trainable parameters, under 3% of network capacity) that combines a per-sector scan-difference channel with a two-head self-attention mechanism and a 384→48 compression bottleneck. Policies are trained with Soft Actor-Critic across a seven-phase progressive curriculum with bidirectional demotion, scaling from static goal-seeking to fifteen simultaneously moving obstacles at 0.18 m/s, with a Hardware-Calibrated Training Mode applied in the final phase. Seven configurations were evaluated on three zero-shot benchmark arenas over 300 episodes each across three random seeds, and key control variants were validated in 130 valid physical trials on a TurtleBot3 Waffle Pi across two matched corridor scenarios. The simulation and physical evaluations give different orderings, and this dissociation is the paper’s principal result. In simulation, SAC-PV-STAM, SAC-R-PV-STAM and SAC-MLP-FS lie within 4.6 percentage points of one another on two of three benchmarks and are not separable; on hardware, they separate with large margins. SAC-R-PV-STAM reached the goal in 19 of 20 static trials and 20 of 20 trials with a moving obstacle, with no threshold violations across 40 trials, against SAC-MLP-FS (stratified p = 0.00068) and SAC-PV-STAM (stratified p = 1.1 × 10−7). A non-recurrent variant stalled with a clear floor ahead—SAC-PV-STAM commanded no forward velocity above 0.005 m/s in any of the 2746 samples of the final 91.3 s of a 122.3 s stall while the LiDAR reported 3.50 m directly ahead—and a prediction stated before the experiment, that introducing a moving obstacle would remove the stall, was confirmed (timeouts 9/10 to 0/20, p = 7 × 10−7). We further report a scan-difference contamination analysis showing up to 12.0× more spurious spikes during self-rotation on hardware, a width-matched comparison separating the effect of Huber critic loss from that of critic width, and a recurrent LSTM baseline that remains significantly inferior across all three benchmarks and unstable across seeds.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.