MaskVLA: Visual Masking Against Trajectory Overfitting of Vision-Language-Action Model
MaskVLA, a masking-based fine-tuning strategy that randomly masking a small portion of the main camera's visual information leads to the emergence of robust policies, thereby enhancing the model's capability to tackle complex manipulation tasks and improving its generalization performance.
Yuxuan Jiang, Jia-Ying Huang, Ge Wang et al.
· 0 citations