Preprint
Aug 2026
UniJEPA: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling
UniJEPA is presented, a unified JEPA that jointly learns photometric prediction (image-level transformations) and temporal prediction (video-level next-state dynamics) in one shared latent space and shows that the same latent space supports controllable abstraction.
Andriana Lanji, Dawei Liu, Jin Li et al.
· 0 citations