2025
Aligning What Matters: Masked Latent Adaptation for Text-to-Audio-Video Generation
A novel framework that performs selective cross-modal alignment through a learnable masking mechanism, enabling the model to isolate and align only the shared latent components relevant to both modalities is proposed.
Jiyang Zheng, Siqi Pan, Yu Yao et al.
· Neural Information Processin... · 6 citations