Preprint
Aug 2026
V-RAE: Rethinking Video Latent Spaces for Generation
V-RAE, a video representation autoencoder that builds compact generative latents on top of frozen vision foundation model representations, and tFVD, a temporal-coherence diagnostic that correlates more reliably with downstream generation quality are introduced.
Minghui Guo, Shengqiong Wu, Hao Fei
· 0 citations