Skip to content

Efficient Mamba-Centered Global-Local Refinement for Language-Guided Visual Grounding in Robotic Grasping

2026 · IEEE Transactions on Automation Science and Engineering · Vol 23, pp. 15638-15657 · 0 citations · 56 references

Abstract

Enabling robots to understand natural language and locate referred objects for grasping remains a key challenge. Language-guided visual grounding connects visual perception and language understanding. As a fine-grained setting, Referring Image Segmentation (RIS) further provides pixel-level masks, which are particularly useful for precise grasping. However, existing RIS methods still face difficulties in robotic scenarios. Convolution- and transformer-based models are often limited by restricted receptive fields or quadratic complexity. Meanwhile, Mamba-based models provide efficient long-sequence modeling, but can still suffer from information decay and loss of spatial details. To address these issues, we propose MambaGLR, an efficient Mamba-centered hybrid framework for language-guided visual grounding, instantiated on RIS and tailored for robotic grasping. MambaGLR adopts a stage-specific cross-modal design: a global-local fusion module is applied in early high-resolution stages to capture both long-range dependencies and local spatial details, while a detail-guided cross-modal refinement module explicitly introduces cross-attention in later low-resolution stages to compensate for information decay and strengthen vision-language alignment. In addition, we construct RefGrasp, a grasp-oriented RIS dataset, and establish a unified benchmark with OCID-VLG and RoboRefIt to support future research in robotic visual grounding. Extensive experiments show that MambaGLR achieves strong grounding accuracy with a favorable efficiency-accuracy trade-off, while laboratory tabletop robot experiments demonstrate its feasibility under the evaluated service-oriented grasping settings. The dataset is publicly available at https://github.com/xiaozheng-liu/MambaGLR Note to Practitioners—Practical robotic grasping increasingly requires understanding natural language to select a desired object in cluttered scenes. Referring image segmentation can provide pixel-level masks for precise grasp execution, but many RIS models are either too computationally expensive or lose spatial details, limiting deployment on resource-constrained robots. This paper proposes MambaGLR, an efficient Mamba-centered hybrid framework that combines global-local fusion with cross-modal refinement to improve vision-language alignment while maintaining low overhead. We also introduce RefGrasp, a grasp-oriented RIS dataset and benchmark. Experiments and laboratory tabletop robot tests demonstrate improved grounding accuracy and grasping feasibility under controlled conditions, while industrial deployment and long-term reliability remain to be validated. Future work will extend this framework to language-guided end-to-end grasp detection tasks.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.