Skip to content

GFLA: A Grasping Framework With Learning-Based Perception and Analytical Modeling for Single-View Scenes

2026 · IEEE Transactions on robotics · Vol 42, pp. 3141-3159 · 0 citations · 64 references
Computer Science

Abstract

Antipodal grasping from single-view red-green-blue and depth (RGB-D) images is challenged by occlusion and partial observability, making purely analytical inference ill-posed. We present the Grasping Framework with Learning-Based Perception and Analytical Modeling (GFLA), which fuses learning-based perception with analytical modeling. GFLA projects antipodal contacts to the image plane, samples grasp candidates via inverse projection, and ranks them with a force-closure metric. To compensate for the information loss inherent in single-view observations, we introduce two grasping hypothesis-guided modules: 1) a contact projection detection network that localizes graspable regions and predicts antipodal projections on visible surfaces, and 2) a 3-D U-Net-based scene completion network that completes geometry and provides explicit collision cues. On GraspNet-1Billion, GFLA achieves its largest improvement on the novel object set (average precision (AP) 35.88%, an improvement of 7.59%), demonstrating superior generalization to previously unseen object categories while also attaining a competitive overall AP of 57.84% (an improvement of 1.33%). Real-robot experiments in cluttered environments, without domain adaptation or fine-tuning, achieve grasp success rates of 95.42% for single-object scenes and 90.12% for multiobject scenes, demonstrating strong practical robustness.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Iterative Grasp Pose Refinement: A Deep Reinforcement Learning Approach for 2D Vision

A reinforcement learning-based framework for robotic grasp refinement, integrating keypoint-based object representations with a Deep Q-Network (DQN), is proposed, offering a scalable and adaptable solution for contact-rich manipulation tasks.

Amir Arsalan Nematollahi, Shayan Ahmadi, M. T. Masouleh et al. · 0 citations
Preprint Sep 2026

AURORA: Active Uncertainty-Driven Re-Orientation for In-Hand Reconstruction

Observing objects grasped by a robot hand is challenging due to severe visual occlusions. Although in-hand manipulation can expose hidden surfaces, existing approaches often rely on predefined or open-loop reorientation strategies that do not explicitly target under-observed regions. We propose AURORA, an active 3D rec...

Fei-Yu Zhao, Yue-Tong Li, Chen-Xi Xiao · 0 citations
Aug 2026

Model-agnostic pose estimation for enhanced collaborative robot grasping via binocular vision

A novel 7-DoF grasping pose generation framework that integrates sparse attention and null convolution is introduced, which enhances the model’s ability to capture fine-grained features from point clouds, significantly improving the accuracy of parallel gripping pose estimation.

Hui Zhang, Yue Wang, Kang An et al. · 0 citations
Preprint Sep 2026

MS-MEM: Multi-Skill Manipulation-Enhanced Mapping via Uncertainty- and Disturbance-Aware Action Selection

Accurate scene understanding in confined, cluttered spaces such as shelves is essential for service robots, as many everyday tasks require them to locate and retrieve objects reliably. Yet, it remains challenging due to severe occlusions, restricted accessibility, and the need to avoid excessive scene changes. In this...

Yi-Tian Shi, Jesper Mücke, Nils Dengler et al. · 0 citations
Preprint Aug 2026

VISTA: Visually Inferred Spatial ConTact Attention for Contact-Rich Manipulation

Contact-rich manipulation requires precise interaction feedback. While vision-centric imitation learning is prevalent, external visual observations provide indirect and ambiguous cues about contact states, particularly under occlusion or subtle object--gripper interactions; dedicated tactile or force sensors can provid...

Jiaying Chen, Wen-Long Dong, Yan Huang et al. · 0 citations
Preprint Aug 2026

Fast Generative Grasping via Lie Group-Constrained MeanFlow

Grasp synthesis is a core task in robotic manipulation, for which the solution typically forms a multimodal distribution rather than a point estimate. Generative robotic grasping aims to learn this distribution with deep generative models such as diffusion and flow-based approaches. The iterative nature of such generat...

S. T. Bukhari, Yi Wei, Ruiqi Ni et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.