Skip to content
Preprint

Diffusion Policies for Short-Horizon Planning in Robot Crowd Navigation

Aug 2026 · 0 citations · 15 references
Computer Science

TL;DR

Planning Diffusion Policy Optimization is proposed, an offline-to-online reinforcement-learning framework that uses a diffusion policy to generate short-horizon action chunks for crowd navigation and obtains an improved success rate over strong baselines and ablations demonstrate that action chunks are especially important for the modified bounded benchmark.

Abstract

Robot crowd navigation requires safe and efficient decision-making under dense, dynamic, and multimodal human--robot interactions. Existing reinforcement-learning methods typically output a single reactive action at each timestep, which limits their ability to represent diverse short-term avoidance strategies. We propose Planning Diffusion Policy Optimization (PDPO), an offline-to-online reinforcement-learning framework that uses a diffusion policy to generate short-horizon action chunks for crowd navigation. PDPO is first pretrained on collision-avoidance demonstrations and then fine-tuned online with PPO by treating the denoising process as an internal decision process. During execution, the policy generates a five-step action chunk and applies it in a receding-horizon manner. Furthermore, we observe an evaluation artifact in common crowd-navigation benchmarks: without explicit boundary constraints, learned agents may leave the valid domain and bypass dense crowds. To address this, we introduce a setting in which boundary violations are treated as collisions. Experiments show that PDPO obtains an improved success rate over strong baselines, and ablations demonstrate that action chunks are especially important for the modified bounded benchmark.

View source

Similar papers

Aug 2024

LSTP-Nav: Lightweight Spatiotemporal Policy for Map-Free Multi-Agent Navigation With LiDAR

This paper proposes LSTP-Nav, a lightweight, decentralized navigation framework built on LSTP-Net that maps stacked 2D LiDAR observations, goal information, and velocity feedback directly to action and introduces an HS reward to provide smooth, heading-aware safety feedback, and develops PhysReplay-SimLab to improve tr...

Xingrong Diao, Zhi-Qiang Sun, Jian-Wei Peng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

HorizonFlow: Variable-Length Planning for Offline Goal-Conditioned RL

Recent advances in generative planning have made trajectory inpainting a promising approach to offline goal-conditioned reinforcement learning. However, these methods typically specify the planning horizon before generating plan content, even though the appropriate horizon depends on the route itself. A horizon that is...

JunHyeok Oh, Zian Jang, Byung-Jun Lee · 0 citations
Conference Aug 2026

Physics-informed spatio-temporal graph reinforcement learning for safe crowd navigation

Standard distance-based reinforcement learning (RL) policies often fail in dense, mixed-speed crowds because they evaluate pedestrian threats based primarily on spatial proximity, ignoring actual collision timing. We introduce PSTGNav, a physics-informed navigation framework that embeds kinematic tracking into the RL p...

Hang-You Yu, Yang Liu · 0 citations
Preprint Aug 2026

Planner-Conditioned Diffusion for Coordinated Multi-Agent Exploration

A Planner-Conditioned Diffusion Policy (PCDP) is proposed, trained on demonstrations from multiple planner styles with planner identity as an explicit conditioning input, enabling a single shared model to learn a multimodal trajectory distribution and generate diverse, controllable trajectory candidates from the same o...

M. Teo, Jeric Lew, T. Duhan et al. · 0 citations

Temporal coordination aware reinforcement learning for multi-agent UAV navigation in dynamic environments

T-CARE is introduced, a hybrid learning-heuristic multi-UAV coordination framework that integrates zero-shot constrained action selection with priority-aware temporal reservations and achieves 100% success, 0% collision rate, and no observed persistent starvation or deadlock.

Abhudaya Shrivastava, Z. Obradovic · 0 citations
Preprint Sep 2026

DODGER: Safety-Guided Reinforcement Learning for Robot Navigation Among Dynamic Obstacles

Robots operating in human-centered environments must safely navigate among multiple dynamic obstacles to avoid collisions with people and surrounding infrastructure. Control barrier functions (CBFs) provide an effective mechanism for safety filtering, and recent CBF-based reinforcement learning (RL) methods embed such...

Sanghyuk Park, Kwan-Woo Lee, Taekyung Kim et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.