Jul 2026
VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders
VideoRAE is introduced, a representation autoencoder that converts features from a frozen video foundation model into compact, reconstruction-capable latents for video generation, establishing frozen video foundation representations as compact, versatile, and generation-friendly video latents.
Zhihao Xie, Junfeng Wu, Xinting Hu et al.
· arXiv.org · 0 citations