Skip to content
Preprint

Fine-Tuning VLAs with Self-Demonstrated Generative Control for Multi-Task Manipulation

Aug 2026 · 0 citations · 43 references
Computer Science

TL;DR

This paper proposes a self-supervised method that generates online interaction rollouts from the zero-shot VLA as additional training data for finetuning and demonstrates the success of this approach across test sets probing generalization on a real ALOHA robot and a new simulation benchmark in RoboTwin.

Abstract

State-of-the-art vision-language-action (VLA) models such as $\pi_{0.5}$ exhibit strong semantic understanding, instruction following and task behavior. However, when deployed on new robots, even minor mismatches in hardware configuration relative to pretraining can cause severe performance drops. Finetuning the VLA on in-domain expert data from the new embodiment improves performance on the expert task but leads to a loss in its original instruction following and behavioral priors. In this paper, we propose a self-supervised method that generates online interaction rollouts from the zero-shot VLA as additional training data for finetuning. Our experiments show this finetuning scheme yields strong multi-task policies that, on the target robot, (1) inherit prior tasks distilled from the zero-shot model, (2) enable generalist instruction following, while (3) learning new skills from expert data with improved sample efficiency. We demonstrate the success of our approach across test sets probing generalization on a real ALOHA robot and a new simulation benchmark in RoboTwin. Video results are available at https://self-supervised-control.pages.dev/

View source

Similar papers

Preprint Aug 2026

Decoding Task Progress from VLA Representations

The results suggest that VLAs have rich, linearly readable internal representations of semantic quantities like task progress, and that learning to read these signals offers a lightweight, interpretable path toward monitoring deployed visuomotor policies.

Atiksh Bhardwaj, E. W. Duan, Prithwish Dan et al. · 1 citation
#artificial intelligence Preprint Sep 2026

HIL-UMI: Bringing Human-in-the-Loop Post-Training of Vision-Language-Action Models to Universal Manipulation Interface

Large-scale vision-language-action (VLA) models provide powerful priors for robot manipulation, yet adapting them to a specific deployment remains challenging. Supervised fine-tuning (SFT) on task-specific demonstrations provides a step toward deployment, but faces two persistent limitations: static data provide limite...

Zimu Han, Yiming Zeng, Ji-Yao Zhang et al. · 0 citations
Preprint Sep 2026

DistAL: Distance-based Advantage Learning for VLA Fine-Tuning

Vision-language-action models (VLAs) have trans- formed the field of robotic manipulation in recent years by combining the semantic understanding of LLMs with the precise control of flow-matching policies. Advantage conditioning is a recent technique that iteratively improves VLAs by training a value function on deploy...

Reece O'Mahoney, Ioannis Havoutis · 0 citations
Preprint Sep 2026

IMLE-VLA: Fast Single-Step Action Generation for Vision-Language-Action Policies

Vision-language-action (VLA) policies leverage pretrained vision-language backbones to achieve strong cross-task generalization. A leading design couples this backbone with a dedicated continuous action head trained via diffusion or flow matching. However, such heads rely on iterative multi-step sampling, for example 1...

Kian Hosseinkhani, Qin-He Peng, George Shramko et al. · 0 citations
#machine learning Preprint Sep 2026

Reinforcement Learning for Real-Time Vision-Language-Action Policies

Reinforcement learning fine-tuning on top of large, pretrained Vision-Language-Action (VLA) models offers promise for highly reliable robot deployment. However, because of their scale, modern VLA models suffer from high inference latency, so the observation used to select an action is often stale by execution time, cre...

Perry Dong, Kuo-Han Hung, D. Sadigh et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.