Assembly, wear, and component replacement perturb the sensor extrinsics and joint zeros encoded by a humanoid CAD model. Existing procedures calibrate one sensor pair or require external fiducials. Using only robot-native motion and onboard sensing, we present OmniCalib, a target-free workflow that calibrates the full...
Kai-Xiang Lu, Hai-Yu Lan, Chun-Xia Qiao et al.· 0 citations
Images capture the world at one moment, whereas video reveals how it changes. Image self-supervision learns spatial structure from a single moment, while video methods commonly learn temporal relationships inside a representation computed jointly from several frames. We ask whether watching a scene change can instead i...
Wen Huang, Hang Guo, Jia-Rui Yang et al.· 0 citations
Multi-View Unified Camera Fields (MVUCF), a training-only framework that forms a shared action-facing latent field across views, and real-world humanoid experiments further provide evidence of its practical effectiveness under RGB-only deployment.
Jia-Rui Yang, Ye-Hao Lu, Yu-Ning Su et al.· 1 citation
Overall, RAPolicy substantially outperforms the baselines in aggregate task success while requiring fewer human interventions, demonstrating stable policy improvement and high online training efficiency.
Jia-Rui Yang, Jia-Jin Zhang, Bin Zhu et al.· 0 citations
LAWM-3D is proposed, which introduces three tightly coupled key designs: a multi-view invariant unified action tokenization scheme for learning 3D-aware latent actions, a geometric alignment constraint that anchors intermediate encoder features to a pretrained 3D foundation model, and a non-injective RGB-D joint recons...
Jia-Rui Yang, Jiale Zhange, Jiawei Li et al.· 1 citation
Across three flow-based VLA models on multiple simulated manipulation benchmarks and two real-world tasks, StructRL improves exploration efficiency and OOD performance over prior in-chain baselines, demonstrating the effectiveness of structured action-space exploration for adapting flow-based VLA with RL.
Jia-Rui Yang, Bin Zhu, Jing-Jing Chen et al.· 1 citation
This paper argues that what a VLA needs is not the ability to generate language, but the ability to consume grounded language, and introduces a framework that endows a VLA with language competence through in-context post-training and an agentic tool-use interface.
Jia-Rui Yang, Wen Huang, Jia-Le Zhang et al.· 1 citation
V-Link is proposed, which explicitly recovers visual representations during the vision-language (VL) to action (A) feature transfer and injects them into Action DiT through asymmetric pathways.
Ye-Hao Lu, Jia-Rui Yang, Yu-Ning Su et al.· 0 citations
TemporalFlow-VLA provides a compact, physically grounded interface for exploiting ordered execution history without explicit motion estimation or geometric processing at deployment, and shows its clearest advantage over prior methods on longer-horizon, multi-stage manipulation.
Jia-Rui Yang, Ye-Hao Lu, Yu-Ning Su et al.· 1 citation
OVIP-SG is presented, a unified framework for instance-preserving semantic mapping, functional scene partitioning, and language-guided small, fine-grained object retrieval that outperforms ConceptGraphs under a unified evaluation protocol on Replica.
Tianjing Hao, Hai-Yu Lan, Ang Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.