Planning Diffusion Policy Optimization is proposed, an offline-to-online reinforcement-learning framework that uses a diffusion policy to generate short-horizon action chunks for crowd navigation and obtains an improved success rate over strong baselines and ablations demonstrate that action chunks are especially important for the modified bounded benchmark.
Abstract
Robot crowd navigation requires safe and efficient decision-making under dense, dynamic, and multimodal human--robot interactions. Existing reinforcement-learning methods typically output a single reactive action at each timestep, which limits their ability to represent diverse short-term avoidance strategies. We propose Planning Diffusion Policy Optimization (PDPO), an offline-to-online reinforcement-learning framework that uses a diffusion policy to generate short-horizon action chunks for crowd navigation. PDPO is first pretrained on collision-avoidance demonstrations and then fine-tuned online with PPO by treating the denoising process as an internal decision process. During execution, the policy generates a five-step action chunk and applies it in a receding-horizon manner. Furthermore, we observe an evaluation artifact in common crowd-navigation benchmarks: without explicit boundary constraints, learned agents may leave the valid domain and bypass dense crowds. To address this, we introduce a setting in which boundary violations are treated as collisions. Experiments show that PDPO obtains an improved success rate over strong baselines, and ablations demonstrate that action chunks are especially important for the modified bounded benchmark.
This paper proposes LSTP-Nav, a lightweight, decentralized navigation framework built on LSTP-Net that maps stacked 2D LiDAR observations, goal information, and velocity feedback directly to action and introduces an HS reward to provide smooth, heading-aware safety feedback, and develops PhysReplay-SimLab to improve tr...
Xingrong Diao, Zhi-Qiang Sun, Jian-Wei Peng et al.· IEEE Transactions on Automat...· 0 citations
Recent advances in generative planning have made trajectory inpainting a promising approach to offline goal-conditioned reinforcement learning. However, these methods typically specify the planning horizon before generating plan content, even though the appropriate horizon depends on the route itself. A horizon that is...
Standard distance-based reinforcement learning (RL) policies often fail in dense, mixed-speed crowds because they evaluate pedestrian threats based primarily on spatial proximity, ignoring actual collision timing. We introduce PSTGNav, a physics-informed navigation framework that embeds kinematic tracking into the RL p...
Hang-You Yu, Yang Liu· International Conference on...· 0 citations
A Planner-Conditioned Diffusion Policy (PCDP) is proposed, trained on demonstrations from multiple planner styles with planner identity as an explicit conditioning input, enabling a single shared model to learn a multimodal trajectory distribution and generate diverse, controllable trajectory candidates from the same o...
T-CARE is introduced, a hybrid learning-heuristic multi-UAV coordination framework that integrates zero-shot constrained action selection with priority-aware temporal reservations and achieves 100% success, 0% collision rate, and no observed persistent starvation or deadlock.
Abhudaya Shrivastava, Z. Obradovic· Integrated Computer-Aided En...· 0 citations
Robots operating in human-centered environments must safely navigate among multiple dynamic obstacles to avoid collisions with people and surrounding infrastructure. Control barrier functions (CBFs) provide an effective mechanism for safety filtering, and recent CBF-based reinforcement learning (RL) methods embed such...
Sanghyuk Park, Kwan-Woo Lee, Taekyung Kim et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.