Efficient Mamba-Centered Global-Local Refinement for Language-Guided Visual Grounding in Robotic Grasping
Enabling robots to understand natural language and locate referred objects for grasping remains a key challenge. Language-guided visual grounding connects visual perception and language understanding. As a fine-grained setting, Referring Image Segmentation (RIS) further provides pixel-level masks, which are particularl...