StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation
StageWAM is introduced, which augments a Motus-based World Action Model with Stage-JEPA, a goal-conditioned Joint-Embedding Predictive Architecture (JEPA) predictor, which uses a frozen V-JEPA2 encoder to extract the current-state representation and predicts the latent target of the next inferred stage.