Skip to content
Conference

Breaking the Resource Barrier: Parameter-Efficient Hierarchical VLA Fine-Tuning via Single-View Semantic Reasoning

Jul 2026 · 2026 23rd International Conference on Ubiquitous Robots (UR) · pp. 343-346 · 0 citations · 24 references
Computer Science

Abstract

As Vision-Language-Action (VLA) models continue to scale in the number of parameters, the computational cost and resource requirements for domain-specific fine-tuning have become significant barriers to practical robotic deployment. While Parameter-Efficient Fine-Tuning (PEFT) techniques like LoRA (Low-Rank Adaptation) offer a potential solution, they often fail to match the task success rates of their fully fine-tuned counterparts. In this paper, we propose a novel hierarchical VLA architecture that achieves state-of-the-art performance while maintaining high parameter efficiency. Our model decomposes control into a high-level System 2 for instruction-conditioned semantic context encoding—a frozen PaliGemma-3B backbone with 0.12B trainable LoRA parameters and a reactive System 1 for multimodal fusion and action generation. To optimize training efficiency, System 2 takes only a single egocentric image, while System 1 recovers missing context by integrating wrist-view images via ResNet-34 and proprioceptive state history encoded with a single linear projection layer. This information is fused through a Transformer Encoder, and final action trajectories are refined via a Transformer-parameterized conditional flow-matching decoder. To improve task performance, we generate diverse candidates by sampling from N independently initialized Gaussian noise vectors and using different numbers of denoising steps K per sample, and then select the executed action using a Cal-QL (Calibrated Q-Learning)-based critic. Evaluated on the standardized LIBERO benchmark, our proposed model achieved a 98.1% average success rate, outperforming contemporary fully trained models across all task suites. These results demonstrate that strategic architectural design can enable parameter-efficient models to exceed the performance of full-scale fine-tuning, offering a viable path for high-performance robotics under constrained computational resources.

View source

Similar papers

Preprint Aug 2026

Fine-Tuning VLAs with Self-Demonstrated Generative Control for Multi-Task Manipulation

This paper proposes a self-supervised method that generates online interaction rollouts from the zero-shot VLA as additional training data for finetuning and demonstrates the success of this approach across test sets probing generalization on a real ALOHA robot and a new simulation benchmark in RoboTwin.

Prachi Garg, Steve Xing, Prahit Yaugand et al. · 0 citations
Preprint Aug 2026

RA-VLA: Retrieval-Augmented VLA for Test-Time Adaptation

RA-VLA is presented, a retrieval-augmented VLA framework that integrates behavior-aligned context retrieval with a grounded execution pipeline that facilitates seamless task adaptation while preserving inference efficiency.

Sanghwan Jang, Minjin Jeon, Minsoo Kim et al. · 1 citation
Preprint Aug 2026

Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world. This makes representation quality a central bottleneck: under a fixed robot-d...

Senqiao Yang, Chengyao Wang, Yuxin Chen et al. · 2 citations
Preprint Aug 2026

Hierarchical Skill Retrieval for Data-Efficient Adaptation of Vision-Language-Action Models

While Vision-Language-Action (VLA) models pretrained on large-scale robot datasets provide a strong foundation for robot manipulation, their performance can degrade when adapted to new tasks with limited task-specific demonstrations. Retrieval offers a practical way to reuse existing demonstrations for data-efficient a...

Haoran Hao, Shahram Najam Syed, Jeff G. Schneider et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.