Skip to content
Open access

DexGraspDiffuser: Target-Coupled Grasp and Action Diffusion for Dexterous Grasping

Jul 2026 · Biomimetics · Vol 11, pp. 465 · 0 citations · 56 references
Medicine

TL;DR

Compared with the reproduced UniDexGrasp-T baseline under the same object split and evaluation protocol, DexGraspDiffuser improves the three-split average success rate by 3.3 percentage points and reduces the average mean position error by 0.53 cm, indicating that target-coupled grasp and action diffusion contribute to improved grasp quality, execution accuracy, and closed-loop stability.

Abstract

Dexterous grasping with multi-finger robotic hands is essential for general-purpose robotic manipulation, but remains challenging due to high-dimensional hand configurations, multimodal grasp distributions, and contact-rich execution dynamics. Existing methods often decouple grasp target generation from execution policy learning, which limits the consistency between generated grasp goals and downstream control. To address this problem, we propose DexGraspDiffuser, a target-coupled grasp and action diffusion framework for dexterous grasping. The first stage, GraspDiffusion, generates diverse and physically plausible target grasps from object point clouds using a compact representation of hand root translation, continuous rotation, and finger joint configuration. The second stage, a Goal-Conditioned Diffusion Policy, predicts temporally coherent action sequences conditioned on the selected target grasp and current observation. During inference, receding-horizon execution enables action-prefix execution and online replanning for improved robustness. Experiments demonstrate that DexGraspDiffuser achieves success rates of 0.76, 0.72, and 0.68 on training objects, unseen objects from seen categories, and objects from unseen categories, respectively. These results correspond to a three-split average success rate of 0.72 and a train-to-unseen generalization gap of 0.08. Compared with the reproduced UniDexGrasp-T baseline under the same object split and evaluation protocol, DexGraspDiffuser improves the three-split average success rate by 3.3 percentage points and reduces the average mean position error by 0.53 cm. This indicates that target-coupled grasp and action diffusion contribute to improved grasp quality, execution accuracy, and closed-loop stability.

Read PDF

Similar papers

Preprint Jul 2026

GraspGraphNet: Graph-Structured Multi-Embodiment Dexterous Grasp Generation

GraspGraphNet is introduced, a topology-aware grasp generation framework that represents each hand as a URDF-derived kinematic graph and directly generates executable palm poses and joint configurations and suggests that graph-structured hand representations can effectively support dexterous grasp generation across robot hands with different kinematic structures.

Y. Lee, Taeyeop Lee, Hyosup Shin et al. · 0 citations
Preprint Aug 2026

Real-World Cooperative Bimanual Dexterous Grasp of Large Objects from Single-View Observations

This work proposes a real-world bimanual grasping framework that includes a multimodal dataset capturing joint angles, visual observations and force signals; a Denoising Diffusion Probabilistic Model (DDPM)-based module that generates joint-level grasp configurations from segmented point clouds; and an execution strategy that integrates motion planning with online grasp refinement to ensure physical stability and feasibility.

Ziming Li, Mingxuan Wu, Jiaqi Zhang et al. · 0 citations
Preprint Jul 2026

CoorGrasp: Coordinated Contact Control for Adaptive Dexterous Grasping Under Uncertainty

While recent research has focused heavily on dexterous grasp pose generation, less attention has been devoted to the execution of planned grasps. Under shape and position uncertainty, open-loop execution often yields uncoordinated contacts, causing undesired in-hand object motion and even grasp failures. To address this, this paper proposes a tactile-driven model predictive controller for adaptive and delicate execution of diverse dexterous grasps. Our approach emphasizes multi-contact coordination across both approaching and grasping phases, with three key novelties: (i) coordination-aware phase separation, (ii) arm-hand coordination to compensate for position errors, and (iii) adaptive force coordination to increase contact forces in a balanced manner. An analytical model is employed to relate contact forces to robot joint motions for predictive control. Our formulation imposes no restrictions on grasp types or contact configurations and integrates seamlessly with state-of-the-art grasp pose generation methods. We validate the approach through large-scale simulations involving 15k grasps across 478 objects on three robotic hands, and real-world experiments on 8 objects. Results demonstrate that our method achieves higher grasp success rates and reduced undesired object movements.

Mingrui Yu, Yongpeng Jiang, Yongyi Jia et al. · 0 citations
Preprint Aug 2026

DexMani: Human-Derived Manipulability Guidance for Dexterous Rotation

Dexterous object rotation is a sequential contact problem: each support, release, and re-contact decision must both produce the desired object motion, and prepare the hand configuration for continued rotation. Existing reinforcement learning methods discover such movement patterns through trial and error on specific robotic hand embodiments, without explicitly accounting for how each contact transition affects the hand's ability to sustain object rotation in subsequent steps. We introduce DexMani, a framework that transfers human demonstrations as contact-conditioned manipulability evolution. This prior captures how successful human contact transitions reshape the object-rotation directions available to the hand. DexMani then learns this manipulability evolution and uses it to guide downstream reinforcement learning, enabling rotation skills to be acquired across robot embodiments with distinct kinematics and active-contact configurations. Across the Shadow Hand, Allegro Hand, and XHand, DexMani achieves the highest success rates in every evaluated setting for both seen and unseen objects. DexMani reaches an average success rate of 57.5% on LEAP Hand, outperforming other baselines and producing smoother rotatory motions. Project site: https://dexmani.github.io

Xiaoyang Chen, Shengcheng Luo, Haoran Guo et al. · 0 citations
Preprint Aug 2026

CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning

CoToGrasp is a novel generative framework that synthesizes diverse, stable grasps strictly conditioned on specific contact topologies, and introduces a feature-based canonical workspace that projects local object features into a unified gripper-centric domain, effectively decoupling the semantic functional intent from the arbitrary object geometry.

Julien Mérand, Boris Meden, Liming Chen et al. · 1 citation
Preprint Aug 2026

GOAG: Generative and Object-Agnostic Grasp Planner for Dexterous Robotic Manipulation

Multifingered grasping is a crucial robotic skill, but current deep-learning grasp planners often struggle to generalize to new objects because they are trained on limited, object-specific datasets. We introduce a fundamentally different approach, grounded in the observation that the gripper and the object share identical surface geometry at their mutual contact points. We propose GOAG: Generative and Object-Agnostic Grasp Planner for Dexterous Robotic Manipulation, a novel deep generative model that learns a compact latent representation of a specific gripper's contact surface distribution, enabling the efficient sampling of valid grasp configurations without relying on object-specific training data. We show that by introducing object features only at inference time, our model can effectively retrieve admissible contact areas that are compatible with the gripper's capabilities. We validate our approach through extensive experiments on established grasp protocols in both simulated and real-world scenarios, demonstrating its effectiveness with different grippers from the literature. Our method delivers state-of-the-art results on the objects from the MultiDex dataset, achieving an average success rate of 86.93%. It offers significantly faster processing when generating numerous grasps, while matching the performance of leading approaches specifically trained on this dataset. Unlike these methods, our approach does not rely on object-specific training data, highlighting the advantages of object-agnostic learning. It effectively addresses the generalization challenges faced by traditional data-driven grasp planners. Code and videos are available on our project website https://cea-list.github.io/goagweb/ .

Julien Mérand, Boris Meden, Mathieu Grossard et al. · 1 citation