Skip to content
Open access

Dual-Arm Manipulation Policy for Target-Cognitive Generalization

Aug 2026 · Electronics · 0 citations · 5 references

TL;DR

TCG-BP (Target-Cognitive Generalization Bimanual Policy), a target-prior-driven bimanual manipulation policy that converts language target descriptions into temporally consistent pixel-level target masks, and enhances visual representations through image–mask collaborative encoding and fusion is proposed.

Abstract

Robotic manipulation policies have made significant progress in recent years, yet their target-cognitive generalization capability remains insufficient when facing unseen targets and scenarios with similar distractors. Existing methods mostly rely on implicit alignment between language descriptions and global visual features. When target appearance or geometric shape changes, or when similar distractors are present, they struggle to stably establish the correspondence between the language-specified target and action generation, thereby affecting manipulation success rates. To address this problem, this paper proposes TCG-BP (Target-Cognitive Generalization Bimanual Policy), a target-prior-driven bimanual manipulation policy. The method converts language target descriptions into temporally consistent pixel-level target masks, and enhances visual representations through image–mask collaborative encoding and fusion. In the action generation stage, the global scene representation and target-focused representation are extracted from the enhanced visual representations and injected into the policy network in a differentiated manner, enabling continuous action prediction to be constrained by scene context while being guided by target priors. On the RoboTwin 2.0 benchmark, TCG-BP improves the average success rate over π0 by 10.2, 12.2, and 13.8 percentage points under the Seen, Unseen Object, and Unseen Distractor settings, respectively. Experimental results verify the effectiveness of the proposed method.

Read PDF

Similar papers

Preprint Sep 2026

HINT: Human-Intent Inception for Long-Horizon Robot Manipulation

Humans can perform complex manipulations given a simple intent through an overall instruction, while continuously adapting to evolving visual observations. However, current vision-language action (VLA) models and other action policies struggle to realize this high-level intelligent behavior under dense, evolving visual...

Ming-Yu Mei, Haojie Xu, Shi-Hao Jin et al. · 0 citations
Preprint Sep 2026

Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

Robot foundation models achieve strong in-distribution performance but often degrade under visual distribution shifts. When learning to generate actions from pretrained visual representations, models may exploit task-irrelevant visual cues that correlate with demonstrated actions within the training distribution. Such...

Jian-Man Lin, S. Shailesh, Zhong-Yi Luo et al. · 0 citations
Open access Aug 2026

Action- and Language-Conditioned Video Assessment for Embodied Control

ALVA (Action- and Language-Conditioned Video Assessment), a trajectory evaluator that conditions its assessment on visual observations, the executed action sequence, and the natural language instruction, provides more effective feedback than the evaluated static image and embedding-based visual baselines and reduces th...

Hwanhee Kim, Jaehyun Jang, Seung-Min Cha et al. · 0 citations
Jul 2026

Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation

Vision-and-Language Navigation in continuous environments (VLN-CE) requires an agent to ground language in egocentric observations and plan in unseen scenes. Although recent multimodal large models and world-model-based methods have improved navigation, they often preserve excessive task-irrelevant detail, weakening ge...

Yihao Wu, Chen-Yi Xu, Li-Qi Yan et al. · 0 citations
Preprint Sep 2026

SAVLA: Symmetry-Aware Vision-Language-Action Models for Robotic Manipulation

Vision-language-action (VLA) models have become the dominant paradigm for language-conditioned robot manipulation. However, although images and language instructions inherently encode geometric information, VLAs acquire their spatial competence purely from demonstrations. As a result, they are reliable only within the...

Jun-Le Li, Weixian Waylon Li, Fu-Xiang Wu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.