It is suggested that LoRA at r=32 with full vision encoder fine-tuning is a practical approach, reducing static peak VRAM from 36.2 to 10.8 GiB (parameters and optimizer states, activation memory excluded) without detectable performance loss.
Abstract
Deploying billion-parameter Vision-Language-Action (VLA) models on industrial hardware requires fine-tuning to bridge the embodiment gap. Full Fine-Tuning (FFT) provides maximal plasticity but requires data centre-grade GPUs. We present a systematic study of Low-Rank Adaptation (LoRA) for $\pi_0$, a flow-matching VLA, evaluated on four precision assembly tasks with a UR5e robotic manipulator. Across a sweep of LoRA ranks (r=8 to 256), allocation strategies, and component-freezing ablations, we find no statistically significant advantage of FFT over certain LoRA configurations. Performance saturates at r=32, and uniform allocation across the Vision-Language-Model (VLM) backbone and action expert proves sufficient. Freezing the VLM or restricting the vision encoder to LoRA significantly degrades performance, indicating that embodiment adaptation requires both semantic and visual plasticity. These results suggest that LoRA at r=32 with full vision encoder fine-tuning is a practical approach, reducing static peak VRAM from 36.2 to 10.8 GiB (parameters and optimizer states, activation memory excluded) without detectable performance loss.
NebulaVLA is presented, an asynchronous dual-frequency architecture that decouples high-level semantic reasoning from low-level action control, optimizing computational resources and modularity and introduces GESTURE-7, a unified language-grounded action representation.
Congyu Zhao, Shuai Tian, Xu Zhang et al.· 0 citations
This paper proposes a self-supervised method that generates online interaction rollouts from the zero-shot VLA as additional training data for finetuning and demonstrates the success of this approach across test sets probing generalization on a real ALOHA robot and a new simulation benchmark in RoboTwin.
Prachi Garg, Steve Xing, Prahit Yaugand et al.· 0 citations
As Vision-Language-Action (VLA) models continue to scale in the number of parameters, the computational cost and resource requirements for domain-specific fine-tuning have become significant barriers to practical robotic deployment. While Parameter-Efficient Fine-Tuning (PEFT) techniques like LoRA (Low-Rank Adaptation)...
Woo-Kyoung Jeong, Yongwoo Gu, June-sup Yi et al.· 2026 23rd International Conf...· 0 citations
A cache-efficient lifelong Vision-Language-Action learning framework for robotic manipulation, which alleviates the plasticity-stability trade-off with a dual-timescale adaptation mechanism while achieving low-cost robotic deployment with a cache-efficient replay strategy.
Yao He, Gan Sun, Wen-Qi Liang et al.· arXiv.org· 2 citations
Xiao-Robotics-1 serves as a strong robot foundation policy that can be efficiently fine-tuned on complex, dexterous tasks with high data efficiency and across multiple simulation benchmarks, Xiaomi-Robotics-1 outperforms state-of-the-art methods.
Xiaomin Guo, Piao-Piao Jin, Jason Li et al.· arXiv.org· 16 citations· ⚡2
A mechanism-by-mechanism transition to the recurrent behavior-cloning baseline shows that replacing the predictive pathway with direct observation input produces the largest single performance drop, accounting for approximately $70\% of the endpoint gap.
Hiroki Sawada, Shunichi Kasahara· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.