Skip to content

Decoupled Early Exits for Task-Dependent Compute Allocation in Flow-Matching VLAs

Sep 2026 · 0 citations · 27 references
Computer Science

TL;DR

This work proposes a framework that exposes backbone depth V, action expert depth A, and denoising steps $D$ as three jointly configurable compute axes in a VLA, and introduces a KV Cache synthesis mechanism that manages the missing keys and values of the skipped backbone layers, allowing the action expert to exit deeper than the backbone.

Abstract

Flow-matching Vision-Language-Action (VLA) models have emerged as a potential solution for generalist robot control, designed by combining a pretrained Vision-Language Model (VLM) backbone with an action expert that generates continuous robot actions. While these models exhibit impressive capabilities, due to their very high number of parameters, their computational requirements are often prohibitive for robotics control. To mitigate these inefficiencies, existing methods predominantly skip VLM backbone layers with early exits or reduce denoising steps, while leaving action expert depth untouched. We propose a framework that exposes backbone depth $V$, action expert depth $A$, and denoising steps $D$ as three jointly configurable compute axes in a VLA. Starting from a pretrained VLA, we attach lightweight Exit Transformers (ET) at intermediate depths in both the backbone and the action expert, trained to distil the last layer of the policy into each exit. Furthermore, we introduce a KV Cache synthesis mechanism that manages the missing keys and values of the skipped backbone layers, allowing the action expert to exit deeper than the backbone. Finally, we show that the optimal compute budget is task-dependent, with different tasks benefiting from different axes and depths. Notably, our method does not require training the original policy from scratch, and for each exit, it increases the number of parameters by only $2.1\%$ for SmolVLA and $4.1\%$ for $\pi_{0.5}$. We validate our approach across two flow-matching VLAs (SmolVLA, $\pi_{0.5}$) and two benchmarks (LIBERO, Meta-World), revealing complementary effects: $V$ and $A$ respectively reduce FLOPs and latency, while $D$ improves both. Our joint configurations $(V,A,D)$ reduce latency by $79.2\%$ and computation (FLOPs) by $31.8\%$, while improving mean success rate by $5.6\%$.

View source

Similar papers

Preprint Sep 2026

IMLE-VLA: Fast Single-Step Action Generation for Vision-Language-Action Policies

Vision-language-action (VLA) policies leverage pretrained vision-language backbones to achieve strong cross-task generalization. A leading design couples this backbone with a dedicated continuous action head trained via diffusion or flow matching. However, such heads rely on iterative multi-step sampling, for example 1...

Kian Hosseinkhani, Qin-He Peng, George Shramko et al. · 0 citations
#machine learning Preprint Sep 2026

Efficient Vision-Language-Action Management and Serving for Robot Factories

Robion is designed, the first VLA serving and management system for multi-robot, multi-model requests on multi-GPU edge servers that meets SLOs and enables flexible model placements on multi-GPU servers, and integrates an intelligent traffic controller.

Dionysios Adamopoulos, Nattapol Chanpaisit, Basel Fakhri et al. · 0 citations
Preprint Sep 2026

Dense to MoE Adaptation for Compact Vision Language Action Policies

The results suggest that dense to MoE adaptation with dynamic expert deactivation is a practical direction for reducing active VLA model size without severe performance loss.

Mu-Chun Niu, Shuang Chen, Yu-Zhou Wu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Fewer Steps, Better Actions: Rethinking Flow-Matching Inference for VLA Policies

Vision-language-action (VLA) policies based on flow matching generate action chunks through repeated evaluations of an action expert. Increasing the number of integration steps raises inference cost, but does not necessarily improve closed-loop success. We propose Coda, which reallocates part of this integration budget...

Zhi-Peng Tang, Xin-Da Chen, Wei-Ning Rao et al. · 0 citations
Preprint Sep 2026

TaskAnchor: Grounding Task State in Reactive VLAs for Long-Horizon Manipulation

Reactive vision-language-action (VLA) policies suffer from task-state aliasing in long-horizon manipulation, where identical multimodal inputs call for distinct, context-dependent actions. Given that pretrained VLAs already possess rich control primitives to express diverse behaviors, we hypothesize that the execution...

Heng-Yan Liu, Wen-Lve Zhou, Bo Yue et al. · 0 citations
Preprint Sep 2026

RAVEL: Asynchronous Rolling Inference for Flow-Based Vision-Language-Action Models

Flow-based vision-language-action (VLA) models are highly effective for generalist robot manipulation, yet their reliance on computationally expensive VLM encoding and multi-step iterative action generation imposes a significant latency bottleneck. The resulting inference latency makes it difficult for robots to respon...

Yu-Han Chen, Ke Yu, Peng-Fei Liu et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.