A cache-efficient lifelong Vision-Language-Action learning framework for robotic manipulation, which alleviates the plasticity-stability trade-off with a dual-timescale adaptation mechanism while achieving low-cost robotic deployment with a cache-efficient replay strategy.
Abstract
Similar to the natural capabilities of humans to sequentially learn new tasks, robots with Vision-Language-Action (VLA) models should possess lifelong learning ability to learn a new task when deployed in open-world environments. However, most recently proposed lifelong learning models aim to effectively learn the current task (plasticity) or maintain high accuracy on previous tasks (stability), while the plasticity-stability trade-off remains largely unsolved in robotic manipulation models. To address this fundamental challenge, we propose a cache-efficient lifelong Vision-Language-Action learning framework for robotic manipulation (i.e., LifelongVLA), which alleviates the plasticity-stability trade-off with a dual-timescale adaptation mechanism while achieving low-cost robotic deployment with a cache-efficient replay strategy. More concretely, we propose a dual-timescale LoRA gating module to decompose VLA adaptation into two lightweight pathways: a short-term adapter for plasticity and a long-term adapter for stable consolidation. These pathways are integrated via a task-aware gate, enabling explicit control of the plasticity-stability trade-off. In the skill replay phase, a cache-efficient stochastic replay strategy is proposed to preserve more balanced retention signals without full-trajectory storage. Finally, experiments show that LifelongVLA outperforms existing baselines, demonstrating efficient skill expansion, robust retention of learned manipulation behaviors, and reduced reliance on retraining for real-world deployment on an xArm robot.
Reinforcement learning fine-tuning on top of large, pretrained Vision-Language-Action (VLA) models offers promise for highly reliable robot deployment. However, because of their scale, modern VLA models suffer from high inference latency, so the observation used to select an action is often stale by execution time, cre...
Perry Dong, Kuo-Han Hung, D. Sadigh et al.· 0 citations
HyMeS, a hybrid learning framework that leverages the reasoning and memory-management capabilities of coding agents to steer a Markovian VLA for memory-dependent manipulation, is proposed, enabling data-efficient compositional generalization.
Yun-Hao Zhao, Zhen-Yang Ni, Haoyang Chen et al.· 0 citations
Hierarchical Robotic Control (HiRoC) is proposed, a hierarchical post-training framework that decouples high-level task planning from low-level action execution and aligns the executor with planner-generated subgoals before reinforcement learning, mitigating the distribution misalignment between planning and execution.
This work proposes a parameter-efficient approach to fine-tune a pretrained VLM for autonomous navigation using an Imperative Learning paradigm, and introduces a unified end-to-end navigation pipeline for natural-language-driven robotic control.
Sebastian Berger, Katharina Winter, Fabian B. Flohr· 0 citations
LiLa-WAM is proposed, a lightweight world-action model that reasons about the future in a compact latent space and can be trained end-to-end on a single 24GB GPU and the Visual Transition Token (VTT), a language-free task representation that encodes each task as a direction in visual feature space.
Fan Yang, Yu-Ting Su, Xiaobo Wang et al.· 9 citations
Imitation-learned vision--language--action (VLA) foundation models acquire broad manipulation capabilities by scaling robot data across tasks and embodiments, but reliable deployment on a specific downstream task and hardware platform still requires post-training. Dexterous hands make this adaptation particularly diffi...
Jun-Lei Zhu, Shen-Zhe Yao, Chao-Gui Huang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.