Reflex: Enabling Fast and Predictive Vision-Language-Action Models for Reaction-Critical Manipulation
This work proposes ReflexVLA, an efficient VLA model designed for reaction-critical manipulation without large-scale robot-data pretraining, which enhances temporal reasoning through latent future prediction and multi-frame temporal fusion within the vision backbone, while reducing deployment latency through batched vi...