Jul 2026
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition
An optimal transport (OT)-based semantic alignment framework for LLM-AVSR is proposed, which explicitly bridges the modality gap by aligning the acoustic and visual representations with reference to the linguistic embedding space of the LLM before multimodal fusion.
Xu-Gang Lu, Peng Shen, Yu Tsao et al.
· arXiv.org · 0 citations