Vision-language-action (VLA) and world-action models (WAMs) often degrade under out-of-distribution task variations despite retaining partial task capability. To recover such capability, we propose RoboIRS, an inference-time internal representation steering method that uses successful and failed rollouts to train linea...
Jiu-Zhou Lei, Chang Liu, Da-You Li et al.· 0 citations
This systematic study of vision-tactile representations across 12 contact-rich manipulation tasks finds that representations preserving spatial structure and temporal continuity consistently support more accurate prediction and stronger planning performance.
Zhi-Yuan Zhang, Po-Kuang Zhou, Kai-Di Zhang et al.· 0 citations
PHR-VLA introduces a lightweight auxiliary future head that, during training, aligns the VLA's internal representations with latent dynamics extracted from future observations, demonstrating that privileged latent dynamics alignment provides an effective training signal for improving anticipatory reasoning in VLA polic...
Davood Soleymanzadeh, Kai-Di Zhang, Zhi-Yuan Zhang et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.