Generative world-action models (WAMs) jointly generate future video and vehicle actions, while their action branches remain primarily optimized by expert imitation. Yet imitation provides no explicit closed-loop geometric verdict for generated trajectories, making verification important during both training and deploym...
Feng-Cheng Yu, Dhruv Parikh, Jun-Jie Ye et al.· 0 citations
World-action models (WAMs) jointly predict how a scene will evolve and how an agent should act, however joint generation alone does not necessarily impose a shared geometric constraint on these predictions. We present PhysWAM, a unified world-action model for autonomous driving that co-denoises multiview video, metric...
Dhruv Parikh, Feng-Cheng Yu, Quan-Kai Gao et al.· 0 citations
Rolling-WAM is presented, a formulation that distributes joint denoising across successive replanning cycles and delivers a 4.5x steady-state replanning speedup over standard joint WAMs.
Ying Zhou, Jun-Jie Ye, Yi-Qi Zhao et al.· 0 citations
Logic-VLA is introduced, a formal-requirement-aware VLA that conditions on Signal Temporal Logic (STL) specifications supplied at inference time, showing that a single VLA can adapt its behavior to varying formal requirements without requiring a separate policy for each specification.
Celina Shiyu Wang, Yiqi Zhao, Junjie Ye et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.