Vision-language models can coordinate long-horizon robot manipulation, yet successful task reasoning still depends on whether local physical interactions produce the intended effects. We study how repeated interaction can improve this capability without updating the base model. We introduce RoboHarn-Evo, a dual-loop ha...
Shi-Feng Bao, Fan-Ding Huang, Yi-Han Lin et al.· 0 citations
Robot-policy benchmarks increasingly cover diverse tasks and preset out-of-distribution conditions, but typically evaluate complete trajectories from predefined initial states. These evaluations often focus on the initialized scene and the final outcome, while paying less attention to the dynamic interaction process. D...
Yang Li, Chen Zhao, Zhuo-Ran Wang et al.· 0 citations
Experiments show that bottom-up learning yields consistent generalization across diverse degradations, while top-down modulation substantially improves monocular depth estimation and video instance segmentation under severe interference, establishing a principled brain-inspired computational approach for advancing arti...
Yihan Lin, Yuguo Chen, Y. Meng et al.· Nature Sensors· 0 citations
JEPA-WAM, a latent WAM built in a pretrained V-JEPA space, which couples latent transition prediction with continuous action generation through a shared predictor, predicts a spatially structured joint current-future target that captures task-shared visual temporal structure between current and future observations, whi...
Yi-Han Lin, Jiawei He, Shi-Feng Bao et al.· 9 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.