TCG-BP (Target-Cognitive Generalization Bimanual Policy), a target-prior-driven bimanual manipulation policy that converts language target descriptions into temporally consistent pixel-level target masks, and enhances visual representations through image–mask collaborative encoding and fusion is proposed.
Abstract
Robotic manipulation policies have made significant progress in recent years, yet their target-cognitive generalization capability remains insufficient when facing unseen targets and scenarios with similar distractors. Existing methods mostly rely on implicit alignment between language descriptions and global visual features. When target appearance or geometric shape changes, or when similar distractors are present, they struggle to stably establish the correspondence between the language-specified target and action generation, thereby affecting manipulation success rates. To address this problem, this paper proposes TCG-BP (Target-Cognitive Generalization Bimanual Policy), a target-prior-driven bimanual manipulation policy. The method converts language target descriptions into temporally consistent pixel-level target masks, and enhances visual representations through image–mask collaborative encoding and fusion. In the action generation stage, the global scene representation and target-focused representation are extracted from the enhanced visual representations and injected into the policy network in a differentiated manner, enabling continuous action prediction to be constrained by scene context while being guided by target priors. On the RoboTwin 2.0 benchmark, TCG-BP improves the average success rate over π0 by 10.2, 12.2, and 13.8 percentage points under the Seen, Unseen Object, and Unseen Distractor settings, respectively. Experimental results verify the effectiveness of the proposed method.
Humans can perform complex manipulations given a simple intent through an overall instruction, while continuously adapting to evolving visual observations. However, current vision-language action (VLA) models and other action policies struggle to realize this high-level intelligent behavior under dense, evolving visual...
Ming-Yu Mei, Haojie Xu, Shi-Hao Jin et al.· 0 citations
Robot foundation models achieve strong in-distribution performance but often degrade under visual distribution shifts. When learning to generate actions from pretrained visual representations, models may exploit task-irrelevant visual cues that correlate with demonstrated actions within the training distribution. Such...
Jian-Man Lin, S. Shailesh, Zhong-Yi Luo et al.· 0 citations
ALVA (Action- and Language-Conditioned Video Assessment), a trajectory evaluator that conditions its assessment on visual observations, the executed action sequence, and the natural language instruction, provides more effective feedback than the evaluated static image and embedding-based visual baselines and reduces th...
Hwanhee Kim, Jaehyun Jang, Seung-Min Cha et al.· Italian National Conference...· 0 citations
Vision-and-Language Navigation in continuous environments (VLN-CE) requires an agent to ground language in egocentric observations and plan in unseen scenes. Although recent multimodal large models and world-model-based methods have improved navigation, they often preserve excessive task-irrelevant detail, weakening ge...
Yihao Wu, Chen-Yi Xu, Li-Qi Yan et al.· arXiv.org· 0 citations
V-Link is proposed, which explicitly recovers visual representations during the vision-language (VL) to action (A) feature transfer and injects them into Action DiT through asymmetric pathways.
Ye-Hao Lu, Jia-Rui Yang, Yu-Ning Su et al.· 0 citations
Vision-language-action (VLA) models have become the dominant paradigm for language-conditioned robot manipulation. However, although images and language instructions inherently encode geometric information, VLAs acquire their spatial competence purely from demonstrations. As a result, they are reliable only within the...