Preprint
Jul 2026
Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space
This work proposes Flowley, an end-to-end, single-stage training architecture that produces soundtracks by combining visual features with textual prompts, and introduces Progressive Soft-masked Cross-Attention, which embeds audio-visual synchronization directly within its attention mechanism, adding zero additional computational cost compared to standard attention layers.
Thanh V. T. Tran, N. Nguyen, Luong Tran et al.
· 0 citations