Vision-Language-Action (VLA) policies have shown promising performance in language-conditioned robotic manipulation. However, most existing VLA systems rely on conventional perspective cameras with limited fields of view, often missing global scene context and leading to unreliable manipulation under visual occlusions,...
Peng Xu, Hao-Ran Lin, Wan-Jun Jia et al.· 0 citations
Quadruped mobile manipulation requires two tightly coupled capabilities: reaching manipulation-ready configurations and maintaining stable contact throughout articulated-object interaction. However, existing methods often terminate navigation near the target, leaving a gap between reachability and manipulation readines...
Hao-Ran Lin, Mingyu Yang, Pengfei Qi et al.· 0 citations
Universal Domain Adaptive Object Detection (UniDAOD) addresses domain shifts in open-world scenarios without assuming pre-defined shared categories between source and target domains. Despite progress in this field, existing methods rely heavily on thresholding techniques to filter private category samples, which introd...
Yuan-Fan Zheng, Jin-Lin Wu, Wu-Yang Li et al.· International Journal of Com...· 0 citations
Cross-View Geo-Localization (CVGL) with OpenStreetMap (OSM) performs well in structure-rich urban environments but collapses in feature-sparse scenes such as rural roads. To study this failure mode, in this work, we introduce CV-FSS, a benchmark that pairs sequential panoramas from five rural regions with aligned OSM m...
Junwei Zheng, Yunyi Huang, Ruize Dai et al.· 1 citation
Quadruped mobile manipulation requires two tightly coupled capabilities: reaching manipulation-ready configurations and maintaining stable contact throughout articulated-object interaction. However, existing methods often terminate navigation near the target, leaving a gap between reachability and manipulation readines...
Haoran Lin, Ming Yang, Pengfei Qi et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.