Skip to content

Assembling Two Parts in One Hand

Sep 2026 · 1 citation · 51 references
Computer Science

TL;DR

A reinforcement learning formulation is presented to solve in-hand assembly: mating two rigid objects within a single dexterous hand, with no second arm and no fixture, which shows robustness to state-estimation errors caused by oc- clusion.

Abstract

A hallmark of human dexterity is the cooperative use of fingers, where different fingers take on distinct yet coordinated roles to accomplish fine manipu- lation, such as capping a pen with the hand that holds it. We study this finger-level coordination through in-hand assembly: mating two rigid objects within a single dexterous hand, with no second arm and no fixture. We present a reinforcement learning formulation to solve this problem in a unified framework, which is driven by a goal relative pose between the two parts. Finger coordination is shaped by a function-based auxiliary reward and regularized toward a single human reference pose, while domain randomization and a fusion of historical proprioception and object observation confer robustness to occlusion-induced estimation noise. The same recipe solves three different assembly tasks (Bottle, Syringe, and Marker). Trained purely in simulation, the policies transfer zero-shot to hardware with a single camera, demonstrating robustness to state-estimation errors caused by oc- clusion. Our experiments also reveal that in-hand assembly places demands on hand morphology and can serve as a benchmark for modern robotic hand systems. Videos and code are available at https://ltbgbird.github.io/in-hand-assembly-page/.

View source

Similar papers

Preprint Sep 2026

Learning In-Hand Object Reaching to General 6D Poses

POISE (Palm-relative Object reaching In SE(3), a sim-to-real reinforcement learning framework for in-hand 6D object pose reaching, combines diverse stable-grasp initialization, goal- and geometry-conditioned control, an adaptive 6D goal curriculum, and a compact reward scheme for pose reaching and grasp preservation.

Jun-Xiao Lin, Tian-Yue Wu, Jie Yin et al. · 0 citations
Preprint Sep 2026

One Demonstration, Many Objects: Generalizing Manipulation via Local Contact Geometry

DemoMimic (Dexterous Motion Mimic), a policy that manipulates objects by focusing on their geometry local to the contact points to improve sim-to-real consistency, yielding a single real-world policy that transfers across objects of varying shape, scale, mass, and friction wherever local contact structure is preserved.

Satvik Sharma, Samrat Sahoo, Huang Huang et al. · 3 citations · ⚡1
Preprint Sep 2026

ArtManip: Category-Level Articulated In-Hand Manipulation

Category-level in-hand manipulation of articulated objects is a formidable yet underexplored challenge for dexterous robotic hands. This difficulty stems from two core bottlenecks: first, controlling an object's internal degrees of freedom is tightly coupled with maintaining grasp stability on a free-floating base; sec...

Yang Yang, Teng-Yu Liu, Pu-Hao Li et al. · 0 citations
Preprint Sep 2026

The Cartesian Hand: In-Hand Manipulation with All-Linear Fingers

Robotic manipulation has increasingly pursued human-like dexterous hands with many articulated degrees of freedom, offering rich manipulation capabilities at the cost of mechanical and control complexity. At the other extreme, parallel grippers are simple and robust, but provide little ability to manipulate an object a...

Bo-Xi Xia, Bo-Kuan Li, Ryan Shin et al. · 0 citations
Preprint Sep 2026

FINGR: Learning Dexterous Hand Control for Real-World Rubik's Cube Solving

Manipulating a Rubik's Cube with a single dexterous hand is a challenging test of sustained, contact-rich control: the hand must execute successive layer turns while keeping the cube secure. Each turn requires some fingers to support the cube while others push a moving layer, release contact, and reset for the next mov...

Yutong Liang, Quanquan Peng, Matthew Kim et al. · 0 citations
Preprint Sep 2026

A Reconfigurable Dual-Opposition Architecture for Single-Hand Assembly and Manipulation

In-hand assembly is constrained by the need to maintain grasps on two separate parts while controlling their relative motion within a single hand. To enable both in-hand assembly and manipulation, we present a reconfigurable dual-opposition architecture. Specifically, to support simultaneous grasping of two parts and c...

William Su, Yunosuke Nakamura, Yi-Xiao Wang et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.