Skip to content
Conference

Chongqing dialect speech recognition based on fine-tuning lightweight pretrained models

Aug 2026 · International Conference on Artificial Intelligence, Big Data and Electrical Automation · Vol 14319, pp. 143191D - 143191D-7 · 0 citations · 13 references
Engineering

TL;DR

This work provides a lightweight solution for Chongqing dialect speech recognition and a valuable reference for low-resource dialect research.

Abstract

To address the low accuracy and poor adaptation of generic speech recognition models to unique pronunciations and vocabularies in Chongqing dialect scenarios, this paper constructs a multi-scenario and multi-speaker Chongqing dialect speech dataset. Based on the FunASR framework, we fine-tune the lightweight pretrained model SenseVoiceSmall. Four groups of controlled experiments are designed: baseline, SpecAugment augmentation only, domain hotword enhancement only, and their combination. Results show that the joint optimization strategy achieves the best performance: the model’s Average Correctness (Avg Corr) increases from 81.67% to 84.48%, and Average Character Error Rate (Avg CER) decreases from 25.31% to 20.09%. The recognition of colloquial expressions, unique vocabularies, and typical pronunciations of Chongqing dialect is significantly improved. Meanwhile, the model achieves a Real-Time Factor (RTF) of 0.005 with an average inference latency of 0.045 seconds per utterance, demonstrating high efficiency for lightweight deployment. This work provides a lightweight solution for Chongqing dialect speech recognition and a valuable reference for low-resource dialect research.

View source

Similar papers

Preprint Aug 2026

MERaLiON-GR: Speech Gender Recognition Model for English and SEA Languages

We present MERaLiON-GR, a speech gender recognition system that performs binary classification (female / male) on English and Southeast Asian (SEA) languages. The model finetunes MERaLiON-SpeechEncoder-2, a large conformer based transformer pre-trained on a broad speech corpus, and applies parameter efficient fine-tuni...

Qiongqiong Wang, A. Aw, Nancy F. Chen et al. · 0 citations
Conference Open access Aug 2026

Peningkatan Automatic Speech Recognition Dialek Bugis-Makassar Menggunakan Data Ucapan Sintetis

Abstract. Speech in the Bugis-Makassar dialect poses a major challenge for Automatic Speech Recognition (ASR) systems, particularly in detecting clitic particles that are essential to utterance meaning but do not exist in standard Indonesian. This study investigates the effect of synthetic speech data augmentation usin...

Muh Fatwah Fajriansyah Marlang, Herdianti Darwis, Huzain Azis et al. · 0 citations
Open access Aug 2026

Generalization Analysis of YAMNet-DNN Architectures in Deepfake Audio Classification

The rapid advancement of speech synthesis and voice conversion technologies has increased the risk of audio deepfake attacks, necessitating robust and generalizable detection systems. This study proposes a deepfake audio detection framework that leverages pretrained YAMNet embeddings as a feature extractor, combined wi...

Hakam Dzakwan Diash, Dwi Arman Prasetya, Alfan Rizaldy Pratama et al. · 0 citations
Review Sep 2026

Summary of the ChinaVoices Challenge 2026: Data, Tasks, Baseline, and Methods

This paper summarizes the ChinaVoices Challenge 2026, which aims to establish unified task definitions and evaluation conditions for Chinese dialect speech processing and to advance multi-dialect identification and automatic speech recognition. The challenge covers 16 dialect categories and defines two tasks: Chinese M...

Yu-Jie Liao, Bing-Shen Mu, Shui-Yuan Wang et al. · 0 citations
#natural language process... Preprint Sep 2026

BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models

We introduce BuzzASR, a collection of language-specialized fine-tuned Whisper models adapted for automatic speech recognition (ASR) in 102 languages. Large end-to-end Transformer-based ASR models such as Whisper have revolutionized ASR, but most prominent models are highly multilingual. As a result, these models often...

Shivam Singh, Aditya Yadavalli, Catherine Arnett et al. · 0 citations
Open access Jul 2026

A comparative analysis of pretrained Wav2Vec XLSR-53 and Whisper-Small models for automatic speech recognition in the Telugu language

This paper investigates the effectiveness of pre-trained models like Wav2Vec XLSR-53 and Whisper-Small for developing ASR systems for the Telugu language, addressing the challenge of limited data availability and demonstrating satisfactory results even when fine-tuned on a smaller dataset.

J. Pushparaj, Muzaffar Ahmad Dar, Sri Gani Kaarthikeya Kammula et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.