Preprint
Jul 2026
World Action Planner: Generalizable Decision-Making with Action-Conditioned World Models
This work proposes World Action Planner, a robot planning system that leverages the reasoning capabilities of Vision-Language Models (VLMs) and the physical grounding of a multi-task pose-image conditioned world model, significantly outperforming state-of-the-art end-to-end policy models such as VLAs and WAMs.
Xiangcheng Zhang, Yilun Du
· 0 citations