Skip to content

Author

Hisashi Kawai

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition

An optimal transport (OT)-based semantic alignment framework for LLM-AVSR is proposed, which explicitly bridges the modality gap by aligning the acoustic and visual representations with reference to the linguistic embedding space of the LLM before multimodal fusion.

Xu-Gang Lu, Peng Shen, Yu Tsao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.