Jul 2026· International Conference on Signal Processing and Communications· pp. 1-5· 0 citations· 23 references
Abstract
Fine-tuning of large pre-trained models, such as Whisper, has gained prominence in the area of speech processing. Since full fine-tuning requires a large amount of data as well as high-ended computational resources, parameter efficient fine-tuning (PEFT) has been the preferred choice among researchers. During the past few years, several PEFT techniques have been developed and have been observed to be extremely effective. Motivated by the success of PEFT, in this paper, we have investigated and documented the efficacy of Low-Rank Hadamard Product Adaptation (LoHA) of Whisper models for children's automatic speech recognition (ASR) task especially in limited data scenario. We have also compared LoHA with a few other PEFT variants. Even though LoHA has been explored for signal processing and federated learning tasks, its impact on children's ASR has not yet been studied. Experimental results presented in this paper indicate that applying LoHA significantly influences ASR performance, achieving a word error rate of 3.0%.
A preliminary study on the adaptation of Whisper for Automatic Speech Recognition in Baniwa, an indigenous Arawakan language spoken in Brazil, Colombia, and Venezuela, demonstrating that multilingual foundation models can be successfully adapted to extremely low-resource indigenous languages.
Leonardo Duart, T. Fonseca, T. Chacon· 0 citations
We introduce BuzzASR, a collection of language-specialized fine-tuned Whisper models adapted for automatic speech recognition (ASR) in 102 languages. Large end-to-end Transformer-based ASR models such as Whisper have revolutionized ASR, but most prominent models are highly multilingual. As a result, these models often...
Shivam Singh, Aditya Yadavalli, Catherine Arnett et al.· 0 citations
LoRA-MoE, a parameter-efficient adaptation method that combines Low-Rank Adaptation with a mixture of experts (MoE) to improve speech recognition for individuals with Parkinson’s disease (PD), demonstrates consistent improvements and stable performance across all severity levels, and its performance is robust to the nu...
Seojin Yoon, Ryul Kim, Sang-Min Lee· IEEE Access· 0 citations
Temporal envelopes carry cues critical to speech intelligibility, yet ASR frontends based on log-mel spectrograms do not explicitly model continuous sub-band envelope structure. This limitation is particularly acute for children's speech, where high acoustic variability demands robust feature representations. We propos...
Edem Ahadzi, Ruchi Pandey, Tomi H. Kinnunen· 0 citations
Background In background noise, speakers adapt their speech production, giving rise to Lombard speech, which often improves speech intelligibility (SI). While intelligibility benefits of Lombard speech have been extensively studied in non-tonal languages, it remains unclear whether spectro-temporal cues, which are crit...
M. Scharf, A. Warzybok, L. Wong et al.· PLoS ONE· 0 citations
The lack of specific standards for evaluating the significant variation in children’s speech and the limited availability of annotated speech samples make it hard to automatically identify misarticulation in children speaking low-resource Indian languages. While techniques for assessing pronunciation using Automatic Sp...
Anushri Ghadge, S. Mahajan, Atharva Deshpande· IEEE Access· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.