ViDAL: A Visual Dynamics-Grounded Action Latent Space for Vision-Language-Action Models
Vision-Language-Action (VLA) models have become a central paradigm for robot policy learning, which predict actions in three forms: raw action chunks, discrete action tokens, or continuous action latents. However, existing action representations primarily model action trajectories, with limited consideration of the vis...