Aug 2026· International Conference on Artificial Intelligence, Big Data and Electrical Automation· Vol 14319, pp. 143191D - 143191D-7· 0 citations· 13 references
Engineering
TL;DR
This work provides a lightweight solution for Chongqing dialect speech recognition and a valuable reference for low-resource dialect research.
Abstract
To address the low accuracy and poor adaptation of generic speech recognition models to unique pronunciations and vocabularies in Chongqing dialect scenarios, this paper constructs a multi-scenario and multi-speaker Chongqing dialect speech dataset. Based on the FunASR framework, we fine-tune the lightweight pretrained model SenseVoiceSmall. Four groups of controlled experiments are designed: baseline, SpecAugment augmentation only, domain hotword enhancement only, and their combination. Results show that the joint optimization strategy achieves the best performance: the model’s Average Correctness (Avg Corr) increases from 81.67% to 84.48%, and Average Character Error Rate (Avg CER) decreases from 25.31% to 20.09%. The recognition of colloquial expressions, unique vocabularies, and typical pronunciations of Chongqing dialect is significantly improved. Meanwhile, the model achieves a Real-Time Factor (RTF) of 0.005 with an average inference latency of 0.045 seconds per utterance, demonstrating high efficiency for lightweight deployment. This work provides a lightweight solution for Chongqing dialect speech recognition and a valuable reference for low-resource dialect research.
We present MERaLiON-GR, a speech gender recognition system that performs binary classification (female / male) on English and Southeast Asian (SEA) languages. The model finetunes MERaLiON-SpeechEncoder-2, a large conformer based transformer pre-trained on a broad speech corpus, and applies parameter efficient fine-tuni...
Qiongqiong Wang, A. Aw, Nancy F. Chen et al.· 0 citations
Abstract. Speech in the Bugis-Makassar dialect poses a major challenge for Automatic Speech Recognition (ASR) systems, particularly in detecting clitic particles that are essential to utterance meaning but do not exist in standard Indonesian. This study investigates the effect of synthetic speech data augmentation usin...
Muh Fatwah Fajriansyah Marlang, Herdianti Darwis, Huzain Azis et al.· Bandung Conference Series St...· 0 citations
The rapid advancement of speech synthesis and voice conversion technologies has increased the risk of audio deepfake attacks, necessitating robust and generalizable detection systems. This study proposes a deepfake audio detection framework that leverages pretrained YAMNet embeddings as a feature extractor, combined wi...
Hakam Dzakwan Diash, Dwi Arman Prasetya, Alfan Rizaldy Pratama et al.· International Journal of Adv...· 0 citations
This paper summarizes the ChinaVoices Challenge 2026, which aims to establish unified task definitions and evaluation conditions for Chinese dialect speech processing and to advance multi-dialect identification and automatic speech recognition. The challenge covers 16 dialect categories and defines two tasks: Chinese M...
Yu-Jie Liao, Bing-Shen Mu, Shui-Yuan Wang et al.· 0 citations
We introduce BuzzASR, a collection of language-specialized fine-tuned Whisper models adapted for automatic speech recognition (ASR) in 102 languages. Large end-to-end Transformer-based ASR models such as Whisper have revolutionized ASR, but most prominent models are highly multilingual. As a result, these models often...
Shivam Singh, Aditya Yadavalli, Catherine Arnett et al.· 0 citations
This paper investigates the effectiveness of pre-trained models like Wav2Vec XLSR-53 and Whisper-Small for developing ASR systems for the Telugu language, addressing the challenge of limited data availability and demonstrating satisfactory results even when fine-tuned on a smaller dataset.
J. Pushparaj, Muzaffar Ahmad Dar, Sri Gani Kaarthikeya Kammula et al.· Frontiers in Artificial Inte...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.