2026· IEEE Transactions on Automation Science and Engineering· Vol 23, pp. 15638-15657· 0 citations· 56 references
Abstract
Enabling robots to understand natural language and locate referred objects for grasping remains a key challenge. Language-guided visual grounding connects visual perception and language understanding. As a fine-grained setting, Referring Image Segmentation (RIS) further provides pixel-level masks, which are particularly useful for precise grasping. However, existing RIS methods still face difficulties in robotic scenarios. Convolution- and transformer-based models are often limited by restricted receptive fields or quadratic complexity. Meanwhile, Mamba-based models provide efficient long-sequence modeling, but can still suffer from information decay and loss of spatial details. To address these issues, we propose MambaGLR, an efficient Mamba-centered hybrid framework for language-guided visual grounding, instantiated on RIS and tailored for robotic grasping. MambaGLR adopts a stage-specific cross-modal design: a global-local fusion module is applied in early high-resolution stages to capture both long-range dependencies and local spatial details, while a detail-guided cross-modal refinement module explicitly introduces cross-attention in later low-resolution stages to compensate for information decay and strengthen vision-language alignment. In addition, we construct RefGrasp, a grasp-oriented RIS dataset, and establish a unified benchmark with OCID-VLG and RoboRefIt to support future research in robotic visual grounding. Extensive experiments show that MambaGLR achieves strong grounding accuracy with a favorable efficiency-accuracy trade-off, while laboratory tabletop robot experiments demonstrate its feasibility under the evaluated service-oriented grasping settings. The dataset is publicly available at https://github.com/xiaozheng-liu/MambaGLR Note to Practitioners—Practical robotic grasping increasingly requires understanding natural language to select a desired object in cluttered scenes. Referring image segmentation can provide pixel-level masks for precise grasp execution, but many RIS models are either too computationally expensive or lose spatial details, limiting deployment on resource-constrained robots. This paper proposes MambaGLR, an efficient Mamba-centered hybrid framework that combines global-local fusion with cross-modal refinement to improve vision-language alignment while maintaining low overhead. We also introduce RefGrasp, a grasp-oriented RIS dataset and benchmark. Experiments and laboratory tabletop robot tests demonstrate improved grounding accuracy and grasping feasibility under controlled conditions, while industrial deployment and long-term reliability remain to be validated. Future work will extend this framework to language-guided end-to-end grasp detection tasks.
General-purpose robot control requires models to understand task intent, identify where to interact, capture how the scene evolves, and generate precise actions. Vision-language-action models provide strong semantic priors but typically do not explicitly model scene dynamics, while world-action models couple visual pre...
Hao-Ran Wen, Wen-Fu Wang, Kun-Song Shi et al.· 1 citation
Long-horizon robotic manipulation requires a policy to bridge task-level semantic reasoning with metric three-dimensional interaction geometry. Existing vision–language–action policies usually acquire geometry implicitly from visual tokens or introduce deterministic intermediate variables only in the image plane, which...
Li Lin, Ming-Hao Shi, Teng-Long Wang· Applied Informatics· 0 citations
Language-Driven Grasp Detection aims to generate executable robotic grasp rectangles based on user instructions described in natural language and visual inputs. User commands typically specify not only the target object but also fine-grained constraints such as specific parts or functional regions. Most existing method...
Qiong Wang, Yang Wang, Zhouchao Fu et al.· IEEE Robotics and Automation...· 0 citations
This work proposes a parameter-efficient approach to fine-tune a pretrained VLM for autonomous navigation using an Imperative Learning paradigm, and introduces a unified end-to-end navigation pipeline for natural-language-driven robotic control.
Sebastian Berger, Katharina Winter, Fabian B. Flohr· 0 citations
Vision-language-action (VLA) models have become the dominant paradigm for language-conditioned robot manipulation. However, although images and language instructions inherently encode geometric information, VLAs acquire their spatial competence purely from demonstrations. As a result, they are reliable only within the...
MaskVLA, a masking-based fine-tuning strategy that randomly masking a small portion of the main camera's visual information leads to the emergence of robust policies, thereby enhancing the model's capability to tackle complex manipulation tasks and improving its generalization performance.
Yuxuan Jiang, Jia-Ying Huang, Ge Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.