Skip to content
Preprint

Technical Report for MERL's Real-TSE Challenge Submission

Jul 2026 · 1 citation · 37 references
Engineering

TL;DR

MERL's submission to the Real-TSE Challenge achieved first place in the second track, demonstrating the critical importance of high-quality data preparation and observing that DNSMOS and speaker similarity are susceptible to over-optimization.

Abstract

Target speech extraction (TSE) has largely been dominated by neural network-based approaches trained and evaluated on synthetic fully overlapped data. The Real-TSE Challenge aims to advance performance on real-world far-field noisy and reverberant recordings. This technical report describes MERL's submission to the Real-TSE Challenge. Rather than proposing a novel model architecture, we built upon the baseline model and focused primarily on data preparation and cleaning. Our system was trained in four stages, beginning with pre-training on fully overlapped mixtures and simulated multi-talker conversations with noise and reverberation applied to both the mixture and the enrollment utterances. We then adapted the model to real-world conditions using noisy far-field recordings with pseudo-targets derived from processed close-talk microphone signals. Our submission achieved first place in the second track, demonstrating the critical importance of high-quality data preparation. Furthermore, we observed that DNSMOS and speaker similarity are susceptible to over-optimization, motivating an investigation of their robustness using adversarial attacks. The results show that both metrics can be driven to extreme values without degrading the token error rate or the VAD-based F1 score.

View source

Similar papers

Preprint Sep 2026

RemixIT-TSE: Progressive Synthetic-to-Real Adaptation for Target Speech Extraction via Target-Aware Supervision and Remixing

Target Speech Extraction (TSE) in real-world conversational scenarios suffers from severe performance degradation due to the domain gap between synthetic training data and complex acoustic environments, where signal-level ground truth is typically unavailable. To address this challenge, we make the first attempt to ext...

Yu Wang, Hai-Xin Guan, Shuang Wei et al. · 0 citations
Preprint Aug 2026

The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge: A Benchmark with Natural Mixtures and Degraded Video

Audio-visual speech enhancement (AVSE) uses a target speaker's visible articulation to recover that speaker's voice from overlapped speech. Yet most evaluation protocols rely on synthetic mixtures and reliable video. The ISCSLP 2026 Real-World AVSE Challenge addresses both gaps. Track~1 combines naturally recorded two-...

Kai Li, Wen-Ze Ren, Jun-Jie Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

FFASR: Benchmarking Far-Field Automatic Speech Recognition using High-Fidelity Simulated RIRs

Far-field automatic speech recognition(ASR) degrades under reverberation, noise, and talker motion, yet the benchmarks that drive model selection emphasize close-microphone speech. We present FFASR, a held-out corpus of 15,637 utterances and an open leaderboard spanning nine conditions, each varying a single acoustic f...

Shivam Saini, Eric Bezzam, G. Götz et al. · 0 citations
#natural language process... Preprint Sep 2026

Xiaomi-CocktailASR-1 Technical Report

Recently, large language model (LLM) based ASR models have achieved significant progress, yet they generally lack support for multi-speaker scenarios, where the cocktail party problem remains a critical bottleneck for further advancing ASR. Existing TS-ASR methods, including end-to-end architectures with speaker embedd...

Yi-Ru Zhang, Hang Su, Li-Chun Fan et al. · 1 citation
#natural language process... Preprint Sep 2026

Qwen-Audio-3.0-ASR Technical Report

In recent years, automatic speech recognition (ASR) has witnessed transformative advancements driven by three complementary paradigms: data scaling, model scaling, and deep integration with large language models (LLMs). However, bridging the gap between academic benchmark performance and real-world production utility r...

Chuan-Meng Bian, Da-Ren Chen, Pei-Xin Chen et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.