Skip to content
Open access

BooST: Bridging Semantics and Motions for Efficient Skill Transfer

Aug 2026 · IEEE Robotics and Automation Letters · Vol 11, pp. 11713-11720 · 0 citations · 39 references
Computer Science

TL;DR

BooST is introduced, a two-stage framework that explicitly bridges semantics and motions to satisfy all three desiderata of skill transfer, and achieves superior few-shot adaptation, cross-domain skill transfer, and robustness to dynamic visual distractors, while maintaining a lightweight yet expressive design suitable for real-world deployment.

Abstract

Skill abstraction—the process of learning reusable and temporally extended behaviors—has emerged as a key paradigm for improving sample efficiency and generalization in robot learning. For efficient skill transfer to real robots, learned skills must generalize across tasks and domains, remain robust to visual and dynamic perturbations, and be efficient enough for practical deployment. However, existing methods typically satisfy only a subset of these properties, as they capture either high-level semantic intent (what) or low-level motion dynamics (how). This incomplete skill transfer yields weak priors for policy learning, thereby demanding substantial in-domain data for downstream adaptation. To address these challenges, we introduce BooST, a two-stage framework that explicitly bridges semantics and motions to satisfy all three desiderata. BooST first leverages a cross-modal VQ-VAE to capture both semantic intent and motion dynamics, yielding a unified skill representation. It then distills this representation into a lightweight policy for efficient downstream adaptation to new tasks. Extensive experiments across simulation and real-robot settings demonstrate that BooST achieves superior few-shot adaptation, cross-domain skill transfer, and robustness to dynamic visual distractors, while maintaining a lightweight yet expressive design suitable for real-world deployment.

Read PDF

Similar papers

Preprint Aug 2026

Progressively Learning Heterogeneous Skills in a Unified Latent Space

To prevent the text-to-motion skill from exploiting shortcut pathways instead of learning language semantics, motion intuition distillation is introduced to ground text-to-motion generation in language semantics and a task-guidance module that dynamically adjusts actions based on high-level language instructions is introduced.

Yueyi Zhang, Ming Gong, Linpu He et al. · 0 citations
Preprint Aug 2026

SkillMemo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation

Skill-Based Memory (SkillMemo) framework is proposed that implicitly decomposes long-horizon demonstrations into latent atomic skills and integrates skill-level features into a dynamic episodic memory bank for solving compositional tasks and consistently enhances both DP and VLA backbones.

Changyuan Wang, Chu-Bin Zhang, Zhen-Yu Wu et al. · 0 citations
2026

DualSkill: Unifying Discrete Stability and Continuous Flexibility for Embodied Control

Learning to execute complex, multi-stage tasks requires skill representations that are both compositionally stable and adaptive in execution. Existing hierarchical approaches often face a fundamental trade-off: continuous skills suffer from representational drift due to unconstrained embedding boundaries, while discrete skills exhibit limited expressivity because their deterministic selection cannot capture the multi-modal nuances required for adaptive execution. This tension makes it difficult to achieve reliable composition and adaptive control within a single framework. To address this, we propose DualSkill, a hierarchical framework that learns stable hard skill primitives and builds adaptive soft skills from them. Specifically, DualSkill acquires discrete hard skills via vector quantization with motion-aware distillation, yielding robust and reusable motion primitives that provide structural anchors for skill composition. Conditioned on these primitives, soft skills are modeled as probabilistic continuous mixtures that adapt skill execution while preserving temporal consistency. DualSkill then predicts future skill intentions autoregressively and decodes them into precise low-level actions. We support DualSkill with both theoretical guarantees on its skill representation and extensive experiments across diverse simulation benchmarks and a real-world robotic platform, showing that it outperforms strong baselines and improves generalization. Note to Practitioners—This paper was motivated by the need for robots to execute complex, multi-step tasks in dynamic environments such as homes, warehouses, and factories. In practice, control systems often struggle to balance modular, reusable skills with smooth transitions, leading to unstable or inefficient behavior when task conditions change. Existing approaches typically force a trade-off: either continuous skills that suffer from representational drift or discrete libraries that result in inflexible behavior. This paper presents DualSkill, a hierarchical framework that bridges this gap by decomposing behaviors into stable hard skills for structural reliability and adaptive soft skills for smooth execution. We validate that DualSkill significantly reduces failure rates in complex manipulation tasks on both simulated benchmarks and physical robots. However, the system still relies on structured training data, which may limit its initial deployment in highly unstructured environments. In the future, DualSkill could be applied to mobile robots and human-robot collaboration, further leveraging its flexible and robust framework for real-world tasks.

Ziru Wang, Long Qian, Hao-Wen Sun et al. · 0 citations
Preprint Aug 2026

Hierarchical Skill Retrieval for Data-Efficient Adaptation of Vision-Language-Action Models

While Vision-Language-Action (VLA) models pretrained on large-scale robot datasets provide a strong foundation for robot manipulation, their performance can degrade when adapted to new tasks with limited task-specific demonstrations. Retrieval offers a practical way to reuse existing demonstrations for data-efficient adaptation, but existing methods often rely on visual similarity, state-action representations, or task-level language matching. These approaches may overlook the hierarchical structure of long-horizon manipulation tasks, where complete task matches are rare but reusable skills are often abundant. To address this challenge, we propose Hierarchical Skill Retrieval (HSR), a retrieval framework for data-efficient VLA adaptation. Specifically, HSR first decomposes a target task into candidate skill sequences. It evaluates each plan based on both semantic plausibility and skill reliability estimated from the prior dataset. The selected decomposition is then used for hybrid retrieval. This combines subtask-level language retrieval with behavior-feature reranking to identify demonstrations that are both semantically relevant and compatible with the target task. Finally, we adapt the policy through a two-stage pretraining and finetuning pipeline, which separates general skill acquisition from task-specific adaptation. Experiments on the LIBERO benchmark and several real-world robot manipulation tasks show that HSR improves the average success rate by 10.3% and 21.3% over the strongest baseline, respectively. These results demonstrate the effectiveness of structured skill-level retrieval for data-efficient VLA adaptation. Videos and code are available at https://hoar012.github.io/HSR-Project.

Haoran Hao, Shahram Najam Syed, Jeff G. Schneider et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.