Oct 2026· IEEE Robotics and Automation Letters· Vol 11, pp. 11649-11656· 0 citations· 30 references
Abstract
Language-Driven Grasp Detection aims to generate executable robotic grasp rectangles based on user instructions described in natural language and visual inputs. User commands typically specify not only the target object but also fine-grained constraints such as specific parts or functional regions. Most existing methods implicitly treat language as a whole rather than as structured semantic descriptions, limiting the ability of language to guide precise grasp generation. In this paper, we propose OPAL-Grasp, a novel end-to-end framework that explicitly learns object-level and part-level semantics to strengthen perceptual understanding. We design an Object-Part Guidance Module that uses joint object- and part-level segmentation supervision to learn fine-grained vision–language representations, with the object branch serving as auxiliary supervision and only part-aware features propagated to the grasp prediction module. A Grasp Imbalance-Aware Loss is introduced to reweight sparse yet critical grasp regions, thereby improving optimization stability and detection accuracy. Extensive experiments on large-scale language-driven grasp benchmarks and real-world robotic setups demonstrate that OPAL-Grasp achieves state-of-the-art performance and exhibits strong robustness and generalization across challenging scenes.
Enabling robots to understand natural language and locate referred objects for grasping remains a key challenge. Language-guided visual grounding connects visual perception and language understanding. As a fine-grained setting, Referring Image Segmentation (RIS) further provides pixel-level masks, which are particularl...
Xiaozheng Liu, Ke-Chen Song, Zeng-Lin Xu et al.· IEEE Transactions on Automat...· 0 citations
Robot demonstration generation requires a system to identify where an interaction should occur, plan a feasible motion, and execute the required contact. HiWE connects these decisions through a point-based interface between visual grounding and language-based planning. PointVLM is instruction-tuned to associate task-re...
Guo-Qing Ma, Ming-Qi Yuan, Chen Gao et al.· 0 citations
This work introduces MiGA, a multi-gripper-aware dataset spanning five distinct gripper types across multiple robots with 103,000 demonstrations, and proposes GVLA, which combines a new multi-gripper tokenizer with adapter-based policy routing that improves zero-shot generalization or few-shot adaptation to new objects...
Hanyi Zhang, Zihong Luo, Tianyu Li et al.· 1 citation
Open-vocabulary grasping on a quadruped manipulator requires more than recognizing the target object. The robot must also select a grasp pose that is both consistent with the task semantics and reliable to execute under body motion and viewpoint changes. In this paper, we present VLEG, an embodied vision-language grasp...
Yu-Xing Ji, Fei Meng, Zishang Ji et al.· Journal of Physics, Conferen...· 0 citations
This paper proposes AdaRoboVLG, a task-adaptive Vision-Language-Grasp (VLG) framework that supports generalizable grasp synthesis across different robotic hands. Unlike existing VLG methods that tightly couple foundation models with end-to-end grasp policies, AdaRoboVLG learns an efficient generalizable base policy tha...
Si-Xu Yan, Shi-Kang Wang, Bin-Hua Huang et al.· 1 citation
Category-level in-hand manipulation of articulated objects is a formidable yet underexplored challenge for dexterous robotic hands. This difficulty stems from two core bottlenecks: first, controlling an object's internal degrees of freedom is tightly coupled with maintaining grasp stability on a free-floating base; sec...
Yang Yang, Teng-Yu Liu, Pu-Hao Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.