Skip to content
Open access

Automated Model Selection for Task-Specific RGB–Tactile Fusion in In-Hand Grasp Pose Estimation

2026 · IEEE Access · Vol 14, pp. 125611-125617 · 0 citations · 18 references

TL;DR

This study investigates in-hand pose estimation of a USB stick that is already held within a robotic gripper and provides a controlled task-specific analysis showing that structured model selection can improve multimodal grasp-pose regression within the evaluated setup.

Abstract

Robotic grasping of small, low-feature objects requires highly precise grasp pose estimation to ensure reliable in-gripper alignment. Vision-based approaches (e.g., YOLO-derived detectors) can localize the object but often fail to recover the exact grasp position under occlusion or partial views, while tactile-only methods lack global context; moreover, existing work rarely evaluates, in a task-specific and systematic manner, which fusion configurations are most suitable for this setting. Motivated by this gap, this paper investigates in-hand pose estimation of a USB stick that is already held within a robotic gripper; initial grasp detection is outside the scope of this work. A deep-learning regression model based on YOLOv11 as a feature extractor was developed to estimate the grasp position using inputs from an RGB eye-in-hand camera and a tactile sensor. Within the defined experimental setup, the early RGB–tactile fusion model selected via NAS achieved the lowest summed lateral error of 1.02 mm, compared to tactile-only (1.35 mm) and RGB-only (1.31 mm) models trained under identical protocols. These results reflect performance under the dataset’s image-level partitioning and indicate improved lateral localization within the evaluated setup. The study does not claim a universally optimal fusion strategy, but provides a controlled task-specific analysis showing that structured model selection can improve multimodal grasp-pose regression within the evaluated setup.

Read PDF

Similar papers

Open access 2026

Grasp Pose Estimation of Articulated Objects Based on Semantic and Geometric Feature Fusion

A deep learning-based grasp estimation model designed to enable robotic manipulation with articulated objects that incorporates the attention-based semantic and geometric feature fusion (ASGF) module improved the grasp success rate in the evaluated setting.

Dongwoo Lee, Yeongmin Kim, Seong-Bo Jo et al. · 0 citations
Aug 2026

Model-agnostic pose estimation for enhanced collaborative robot grasping via binocular vision

A novel 7-DoF grasping pose generation framework that integrates sparse attention and null convolution is introduced, which enhances the model’s ability to capture fine-grained features from point clouds, significantly improving the accuracy of parallel gripping pose estimation.

Hui Zhang, Yue Wang, Kang An et al. · 0 citations
Conference Open access Jul 2026

Monocular Vision Based Control Framework for Grasping

A unified monocular vision-based grasping framework that targets both soft and rigid objects within a single control pipeline, using only RGB input and a position-controlled gripper, and is validated in real-world pick-and-place experiments.

Shail V Jadav, Dongheui Lee · 0 citations
Conference Jul 2026

Conditional Flow Matching for Grasp Pose Generation in Multi-Workpiece Scenes

While most robotic research focuses on household tasks such as bus table arrangement and cloth folding, numerous manipulation tasks remain challenging for industrial applications, particularly the grasping and transportation of scattered workpieces. In this paper, we propose SGDIFF, a vision-guided grasping architecture that integrates a pre-trained vision-language model (VLM) with flow matching-based diffusion model. Given an RGB-D image of a tabletop scene, the VLM first detects each workpiece, outputs its bounding box, and assigns a unique ID. A point cloud is then generated for each detected instance. Subsequently, flow matching model iteratively refines the initially noisy gripper pose to a stable and collision-free grasp for each target workpiece. The proposed method eliminates the need for object-specific models and enables efficient multi-object grasping in cluttered industrial environments. The experiments validates the effectiveness of combining semantic understanding with precise pose refinement for robust industrial automation.

Juan Li, Pengxiang You, Qiong Wu et al. · 0 citations
Conference 2026

Two-stage Monocular 6D Pose Estimation for Small Cubic Objects

This paper studies monocular 6D pose estimation of small cubic objects from a single RGB image and proposes a two-stage manipulation- oriented framework, which achieves the strongest overall balance in ADD-S, translation accuracy, rotation stability, and task-oriented usability metrics.

Xinmiao Du · 0 citations