This paper summarizes the ChinaVoices Challenge 2026, which aims to establish unified task definitions and evaluation conditions for Chinese dialect speech processing and to advance multi-dialect identification and automatic speech recognition. The challenge covers 16 dialect categories and defines two tasks: Chinese Multi-Dialect Identification and Chinese Multi-Dialect Automatic Speech Recognition (ASR). It uses approximately 320 hours of speech across the Reference Set, Open Evaluation Set, and Hidden Evaluation Set. The two tasks use the same evaluation audio, and each includes restricted-data and open-data tracks. We describe the task settings, data, evaluation metrics, and Qwen3-ASR-1.7B baseline, and analyze the leaderboard results and submitted systems. In total, 28 teams submit results, 17 provide system reports, and systems from 15 teams pass the compliance review and are included in the analysis. Most eligible systems outperform the baseline, and the official top-three order remains unchanged on the Hidden Evaluation Set for both tasks. Dialect-level results show that categories with higher identification accuracy generally have lower ASR error rates, although the tasks assess related but distinct capabilities. Leading identification systems commonly exploit dialect-discriminative acoustic representations, whereas leading ASR systems emphasize data normalization, augmentation, and auxiliary CTC objectives. These results provide practical guidance for developing and evaluating Chinese multi-dialect speech processing systems.
Yu-Jie Liao, Bing-Shen Mu, Shui-Yuan Wang et al.· 0 citations
Conversational context provides semantic and acoustic cues across turns for automatic speech recognition (ASR), but relying on historical transcripts can propagate recognition errors and discard pronunciation and speaker information. We present a multimodal conversational-context framework for LLM-based ASR that integrates a scenario-controlled data pipeline, scalable multimodal context training, and systematic evaluation. We construct dialogues around entities and their confusable forms and interleave historical user speech with assistant text responses for supervised fine-tuning. We also introduce MM-ContextASR Bench, which evaluates contextual understanding and entity error correction across five scenarios. Experiments with Qwen3-Omni and Step-Audio-2-mini reveal limitations in handling irrelevant and erroneous history and show that our data construction and training improve context utilization, with multimodal context achieving the highest overall entity recall on both models. Further experiments on accent, dialect, and target-speaker ASR demonstrate the value of historical speech. The benchmark data and evaluation code are publicly available at https://github.com/llh666521/MM-ContextASR.
Long-Hao Li, Jian Tang, Yu-Xiang Kong et al.· 0 citations
We introduce the REAL-TSE Challenge, an IEEE SLT 2026 satellite challenge on target speaker extraction~(TSE) from real conversational recordings. Given a multi-speaker mixture and one or more enrollment utterances from a target speaker, participating systems must recover only the target speech. Unlike simulated read-speech benchmarks, REAL-TSE evaluates Mandarin and English recordings that contain natural overlap, reverberation, noise, channel mismatch, and conversational dynamics. The challenge defines two complementary tracks: an Online track for low-latency streaming extraction and an Offline track for full-context processing. Systems are evaluated with Token Error Rate (TER), Speaker Similarity (SpkSim), DNSMOS, and target-speaker activity F1. This overview paper describes the task definition, datasets, baselines, evaluation protocol, submitted systems, condition-wise findings, and lessons for future real-world TSE benchmarks.
Shuai Wang, Zihan Qian, Ke Zhang et al.· arXiv.org· 1 citation
DiaScriber is proposed, an end-to-end multi-speaker diarization and transcription model built on a speech large language model that achieves superior performance over comparison methods across extensive multi-speaker scenario test sets and demonstrates outstanding generalization ability in unseen multi-speaker scenarios.
Bing-Shen Mu, Xian Shi, Xiong Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.