Preprint
Jul 2026
HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving
HyWorldVLA is proposed, a hybrid world-VLA framework that unifies pixel-level supervision and latent representation learning that significantly outperforms both pixel-based and latent-based world model baselines.
Quanfu Yu, Xiangman Wu, Haoqi Xu et al.
· 0 citations