Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semant...
Jia-Yi Chen, Wen-Xuan Song, Jing-Bo Wang et al.· 0 citations
World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit. This makes it...
Wen-Bo Chen, Tian-Fu Li, Hao-Xuan Xu et al.· 0 citations
Vision-Language-Action models are increasingly effective for robotic manipulation, yet most predict actions directly from current observations without explicitly modeling future scene evolution. Recent methods introduce future prediction to improve action generation, but dense future modeling often requires expensive i...
Tian-Fu Li, Hao-Xuan Xu, Wen-Bo Chen et al.· 0 citations
Safe and efficient trajectory planning is essential in autonomous driving. However, existing end-to-end approaches often fall short in both computational efficiency and safety guarantees. Methods based on imitation learning suffer from causal confusion, while rule-based scoring approaches often incur heavy computationa...
Cheng-Lin Chen, Lu-Jia Wang, Xin-Hu Zheng et al.· 0 citations
This work proposes an object--path graph that unifies open-vocabulary semantic reasoning with topological navigation, and introduces a navigation strategy that combines global path planning with local inter-node execution through lightweight node localization and semantic visual servoing, enabling navigation directly o...
Lin-Wei Zheng, Dao-Jie Peng, Bing-Tao Wang et al.· 0 citations
Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existing methods have two main limitations. First, they underuse dynamics priors shared across diverse visual and interaction d...
Jun-Feng Li, Junjie He, Zhi-De Zhong et al.· 1 citation
Active perception is essential for robotic manipulation when fixed viewpoints leave task-relevant information occluded or unobserved. However, enabling vision-language-action (VLA) models to reason across changing viewpoints and actively acquire informative observations remains challenging. We present ActiveScale, a fr...
Shuai-Tao Zhou, Kai-Sheng Pang, Wen-Xuan Song et al.· 1 citation
SimFuse3D is introduced, which preserves the target placement and repairs the associated pseudo object using measured geometry from labeled source scans, and achieves the best performance among the compared adaptation methods on nearly all metrics.
Yong-Chun Lin, Xin-Liang Zhang, Yun Zou et al.· 0 citations
4D-WAM is proposed, a model-agnostic training strategy that injects spatiotemporal knowledge from 3D trajectory fields into WAMs through representation alignment, enabling WAMs to learn trajectory-level spatiotemporal representations.
Lishan Yang, Wen-Xuan Song, Xi Wang et al.· 5 citations· ⚡1
SSMB is presented, a deblur-free, self-supervised keypoint detector for motion-blurred images that requires neither handcrafted detectors nor external pseudo-labels, and introduces the Local Discriminability Enhancement (LDE) module, which restores fine-grained local discriminability after global feature mixing.
Zhenjun Zhao, F. Bellavia, Wen-Ting Wang et al.· 0 citations
MobileWAM surpasses state-of-the-art mobile manipulation policies on ManiSkill-HAB and fine-tunes to a real ARX Lift2 mobile manipulator across diverse tasks with strong generalization.
Ze-Hua Fan, Jun-Jie He, Wen-Xuan Song et al.· 3 citations· ⚡1
DreamTrajectory is presented, a trajectory-guided framework for language-conditioned mobile manipulation that introduces one component for each limitation of existing Vision-Language-Action policies, and jointly predicts an intention-level end-effector trajectory and a whole-body action chunk in a single action expert.
Zheng Yang, Wen-Jie Zhang, Xiang-Yu Chen et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.