Skip to content

Author

Long-Khanh Pham

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space

This work proposes Flowley, an end-to-end, single-stage training architecture that produces soundtracks by combining visual features with textual prompts, and introduces Progressive Soft-masked Cross-Attention, which embeds audio-visual synchronization directly within its attention mechanism, adding zero additional computational cost compared to standard attention layers.

Thanh V. T. Tran, N. Nguyen, Luong Tran et al. · 0 citations