Skip to content
Preprint

Reflex: Enabling Fast and Predictive Vision-Language-Action Models for Reaction-Critical Manipulation

Aug 2026 · 0 citations · 22 references
Computer Science

TL;DR

This work proposes ReflexVLA, an efficient VLA model designed for reaction-critical manipulation without large-scale robot-data pretraining, which enhances temporal reasoning through latent future prediction and multi-frame temporal fusion within the vision backbone, while reducing deployment latency through batched visual encoding and CUDA Graph replay.

Abstract

Vision-Language-Action (VLA) models have recently achieved promising performance in robotic manipulation. However, existing benchmarks mainly evaluate generalization on static manipulation tasks and largely overlook dynamic interaction scenarios. To address this gap, we present ReflexBench, a benchmark for reaction-critical manipulation. ReflexBench contains six dynamic tasks and introduces an evaluation framework that decouples simulator stepping from robot control while supporting configurable latency under synchronous and asynchronous inference. Building upon ReflexBench, we propose ReflexVLA, an efficient VLA model designed for reaction-critical manipulation without large-scale robot-data pretraining. ReflexVLA enhances temporal reasoning through latent future prediction and multi-frame temporal fusion within the vision backbone, while reducing deployment latency through batched visual encoding and CUDA Graph replay. Experiments show that ReflexVLA consistently improves dynamic manipulation performance while maintaining competitive accuracy on standard static manipulation benchmarks, and real-world experiments further demonstrate its effectiveness under practical deployment conditions. Project website: https://reflexvla.github.io

View source

Similar papers

Preprint Sep 2026

Towards Neuro-Symbolic Procedural Reasoning for Long-Horizon Vision-Language-Action Manipulation

Vision-language-action (VLA) models can execute short manipulation skills, but remain brittle in long-horizon procedures requiring persistent task state, dependency-aware reasoning, conditional decisions, and reliable grounding. We investigate a neuro-symbolic framework that combines learned VLA control with explicit t...

Vivek Chavan, Ya-Huan Shi, O. Heimann et al. · 0 citations
Preprint Sep 2026

Decoupling Vision, Language, and Action for Efficient Multi-Task Robot Policies

Vision-Language-Action (VLA) policies commonly run Vision-Language Model (VLM) backbones with billions of parameters at every policy inference, which costs latency and energy. We revisit a decoupled alternative for multi-task manipulation: separate vision and language encoders whose representations condition a compact...

Xia-Tao Sun, Chen Liang, Zi-Yao Zeng et al. · 2 citations
Preprint Sep 2026

Catch Me If You Can: Real-Time Feedback Denoising for Responsive VLAs

Vision-Language-Action (VLA) models have shown strong generalization in robotic manipulation by combining semantic knowledge from pretrained vision-language models with expressive action-generation policies. Diffusion-based action generators are particularly effective for modeling temporally coherent action chunks, but...

Yi-Heng Ji, Xing-Ru Zhou, Luis Sentis et al. · 0 citations
#machine learning Preprint Sep 2026

Reinforcement Learning for Real-Time Vision-Language-Action Policies

Reinforcement learning fine-tuning on top of large, pretrained Vision-Language-Action (VLA) models offers promise for highly reliable robot deployment. However, because of their scale, modern VLA models suffer from high inference latency, so the observation used to select an action is often stale by execution time, cre...

Perry Dong, Kuo-Han Hung, D. Sadigh et al. · 0 citations
Preprint Aug 2026

DREAM: Deployment-Time Demonstration Generation via Real-to-Sim for Scalable Policy Adaptation

DREAM is presented, a framework that generates fine-tuning data for a pretrained VLA from a captured workspace and a language instruction, without requiring a task-specific human demonstration, and whether it can serve as a scalable data-collection system for the deployment workspace.

Makoto Sato, T. Matsushima, Yutaka Matsuo et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.