Skip to content
Preprint

The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge: A Benchmark with Natural Mixtures and Degraded Video

Aug 2026 · 0 citations · 35 references
Engineering Computer Science

Abstract

Audio-visual speech enhancement (AVSE) uses a target speaker's visible articulation to recover that speaker's voice from overlapped speech. Yet most evaluation protocols rely on synthetic mixtures and reliable video. The ISCSLP 2026 Real-World AVSE Challenge addresses both gaps. Track~1 combines naturally recorded two-talker mixtures, which lack a clean reference, with reference-available remixes of the same speakers; Track~2 additionally degrades the target video in five ways and adds 3-m far-field recordings. Sixteen and twelve teams were ranked on speaker-disjoint test data by rank averaging over waveform fidelity, predicted quality, transcription accuracy, and speaker similarity. On Track~1 remixes, the best system reaches 12.7~dB SI-SDR and 0.85 STOI, but natural recordings remain harder: even the lowest CER rises from 9.0\% to 14.7\%. The leading systems use video mainly for speaker attribution rather than signal reconstruction, and the top two lose under 0.5~dB SI-SDR on Track~2. UTMOS and DNSMOS rank systems differently from the other metrics, so no single metric captures target-speech recovery. We release the baselines, evaluator, and official results.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.