AWM-VLA: AlignedWorld Modeling for Efficient and Explainable Vision-Language-Action Policies
Vision-language-action (VLA) models have become a powerful paradigm for generalist robotic manipulation, yet they are often reactive: the policy maps the current observation directly to an action chunk without reasoning about the long-term consequences of its decisions. Prior attempts to endow policies with world model...