Mobile manipulation extends robot interaction beyond a fixed kinematic workspace by making the reachable region itself controllable. This flexibility introduces two central challenges: spatially grounded perception under continuous ego-motion and coordinated control of heterogeneous arm and base actions. Existing appro...
Qi-Wei Liang, Guang-Yu Chen, Shao-Long Zhu et al.· 0 citations
MoPA is presented, a framework that aligns perceptual conditioning with mobility and manipulation while preserving coordination at the action level, and achieves state-of-the-art performance across all three task suites.
Guang-Yu Chen, Qi-Wei Liang, Shao-Long Zhu et al.· 2 citations
GeoLAM, a framework for learning geometry-grounded latent actions from action-free human videos, combines future-frame reconstruction through a frozen geometric feature hierarchy with motion supervision from a training-only 4D geometry teacher and requires neither the geometry teacher nor future-video generation.
Yi-Fan Xie, He-Kun Tian, Jin-Kun Liu et al.· 0 citations
Active exploration and semantic navigation require an embodied agent to build memory from partial observations, predict how the evolution of observed spatial memory may support future motion, and convert that prediction into actionable plans. We present GLAM, a goal-conditioned latent world model trained over global sp...
I-Tak Ieong, Rui-Zhi Feng, Zhao-Yang Lu et al.· 0 citations
Vision-language-action (VLA) policies commonly execute long-horizon mobile manipulation through open-loop action chunks, issuing multiple actions without receiving new high-level visual input. A committed chunk therefore implies how observations should evolve, but accidental deviations can violate this expectation whil...