Audio-visual speech enhancement (AVSE) uses a target speaker's visible articulation to recover that speaker's voice from overlapped speech. Yet most evaluation protocols rely on synthetic mixtures and reliable video. The ISCSLP 2026 Real-World AVSE Challenge addresses both gaps. Track~1 combines naturally recorded two-...
Kai Li, Wen-Ze Ren, Jun-Jie Li et al.· 0 citations
Multi-encoder fusion extends Large Audio-Language Models (LALMs) beyond speech-centric recognition, but selecting encoders via intuition or exhaustive search often introduces redundant representations and inflates an already constrained compute budget. We propose CUES (Correlation-gUided Encoder Selection), a lightweig...
Pei-Jun Liao, Hung-Shin Lee, Wenze Ren et al.· 0 citations
Experiments on the MER2026-EmoPrefer Challenge dataset and the error-augmented dataset demonstrate that EAPO improves emotion preference prediction and enhances the robustness of MLLM judges when evaluating fluent descriptions that conflict with the video's multimodal emotional evidence.
Zilong Huang, Junyi Peng, Junjie Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.