V-Co: A Closer Look at Visual Representation Alignment via Co-Denoising
V-Co, a systematic study of visual co-denoising in a unified JiT-based framework, outperforms the underlying pixel-space diffusion baseline and strong prior pixel-diffusion methods while using fewer training epochs, offering practical guidance for future representation-aligned generative models.