This study investigates in-hand pose estimation of a USB stick that is already held within a robotic gripper and provides a controlled task-specific analysis showing that structured model selection can improve multimodal grasp-pose regression within the evaluated setup.
Abstract
Robotic grasping of small, low-feature objects requires highly precise grasp pose estimation to ensure reliable in-gripper alignment. Vision-based approaches (e.g., YOLO-derived detectors) can localize the object but often fail to recover the exact grasp position under occlusion or partial views, while tactile-only methods lack global context; moreover, existing work rarely evaluates, in a task-specific and systematic manner, which fusion configurations are most suitable for this setting. Motivated by this gap, this paper investigates in-hand pose estimation of a USB stick that is already held within a robotic gripper; initial grasp detection is outside the scope of this work. A deep-learning regression model based on YOLOv11 as a feature extractor was developed to estimate the grasp position using inputs from an RGB eye-in-hand camera and a tactile sensor. Within the defined experimental setup, the early RGB–tactile fusion model selected via NAS achieved the lowest summed lateral error of 1.02 mm, compared to tactile-only (1.35 mm) and RGB-only (1.31 mm) models trained under identical protocols. These results reflect performance under the dataset’s image-level partitioning and indicate improved lateral localization within the evaluated setup. The study does not claim a universally optimal fusion strategy, but provides a controlled task-specific analysis showing that structured model selection can improve multimodal grasp-pose regression within the evaluated setup.
Experiments show that M-VTOP achieves sub-millimeter accuracy under complex geometries, occlusions, and tight tolerances, demonstrating its promise for high-precision robotic manipulation.
M. Oller, Qiyang Qian, Radu Corcodel et al.· 0 citations
A deep learning-based grasp estimation model designed to enable robotic manipulation with articulated objects that incorporates the attention-based semantic and geometric feature fusion (ASGF) module improved the grasp success rate in the evaluated setting.
Dongwoo Lee, Yeongmin Kim, Seong-Bo Jo et al.· IEEE Access· 0 citations
A novel 7-DoF grasping pose generation framework that integrates sparse attention and null convolution is introduced, which enhances the model’s ability to capture fine-grained features from point clouds, significantly improving the accuracy of parallel gripping pose estimation.
Hui Zhang, Yue Wang, Kang An et al.· Signal, Image and Video Proc...· 0 citations
A unified monocular vision-based grasping framework that targets both soft and rigid objects within a single control pipeline, using only RGB input and a position-controlled gripper, and is validated in real-world pick-and-place experiments.
Shail V Jadav, Dongheui Lee· 2026 IEEE/ASME International...· 0 citations
While most robotic research focuses on household tasks such as bus table arrangement and cloth folding, numerous manipulation tasks remain challenging for industrial applications, particularly the grasping and transportation of scattered workpieces. In this paper, we propose SGDIFF, a vision-guided grasping architecture that integrates a pre-trained vision-language model (VLM) with flow matching-based diffusion model. Given an RGB-D image of a tabletop scene, the VLM first detects each workpiece, outputs its bounding box, and assigns a unique ID. A point cloud is then generated for each detected instance. Subsequently, flow matching model iteratively refines the initially noisy gripper pose to a stable and collision-free grasp for each target workpiece. The proposed method eliminates the need for object-specific models and enables efficient multi-object grasping in cluttered industrial environments. The experiments validates the effectiveness of combining semantic understanding with precise pose refinement for robust industrial automation.
Juan Li, Pengxiang You, Qiong Wu et al.· 2026 6th International Confe...· 0 citations
This paper studies monocular 6D pose estimation of small cubic objects from a single RGB image and proposes a two-stage manipulation- oriented framework, which achieves the strongest overall balance in ADD-S, translation accuracy, rotation stability, and task-oriented usability metrics.
Xinmiao Du· Poster Volume 0007 The 2026...· 0 citations