Jun 2026
Z-1: Efficient Reinforcement Learning for Vision-Language-Action Models
The results show that systematic GRPO post-training can substantially improve flow-based VLA policies without additional private demonstrations, and outperforms the published sota models.
Lang Cao, Renhong Chen, Luyi Li et al.
· arXiv.org · 0 citations