Skip to content
Conference

Adapting Whisper Models Using LoHA for Robust Recognition of Children's Speech

Jul 2026 · International Conference on Signal Processing and Communications · pp. 1-5 · 0 citations · 23 references

Abstract

Fine-tuning of large pre-trained models, such as Whisper, has gained prominence in the area of speech processing. Since full fine-tuning requires a large amount of data as well as high-ended computational resources, parameter efficient fine-tuning (PEFT) has been the preferred choice among researchers. During the past few years, several PEFT techniques have been developed and have been observed to be extremely effective. Motivated by the success of PEFT, in this paper, we have investigated and documented the efficacy of Low-Rank Hadamard Product Adaptation (LoHA) of Whisper models for children's automatic speech recognition (ASR) task especially in limited data scenario. We have also compared LoHA with a few other PEFT variants. Even though LoHA has been explored for signal processing and federated learning tasks, its impact on children's ASR has not yet been studied. Experimental results presented in this paper indicate that applying LoHA significantly influences ASR performance, achieving a word error rate of 3.0%.

View source

Similar papers

#small language model Preprint Aug 2026

Fine-Tuning Whisper for Automatic Speech Recognition in Baniwa: A Preliminary Study

A preliminary study on the adaptation of Whisper for Automatic Speech Recognition in Baniwa, an indigenous Arawakan language spoken in Brazil, Colombia, and Venezuela, demonstrating that multilingual foundation models can be successfully adapted to extremely low-resource indigenous languages.

Leonardo Duart, T. Fonseca, T. Chacon · 0 citations
#natural language process... Preprint Sep 2026

BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models

We introduce BuzzASR, a collection of language-specialized fine-tuned Whisper models adapted for automatic speech recognition (ASR) in 102 languages. Large end-to-end Transformer-based ASR models such as Whisper have revolutionized ASR, but most prominent models are highly multilingual. As a result, these models often...

Shivam Singh, Aditya Yadavalli, Catherine Arnett et al. · 0 citations
Open access 2026

LoRA-MoE Fine-Tuning for Improved Speech Recognition in People With Parkinson’s Disease

LoRA-MoE, a parameter-efficient adaptation method that combines Low-Rank Adaptation with a mixture of experts (MoE) to improve speech recognition for individuals with Parkinson’s disease (PD), demonstrates consistent improvements and stable performance across all severity levels, and its performance is robust to the nu...

Seojin Yoon, Ryul Kim, Sang-Min Lee · 0 citations
Preprint Sep 2026

A Temporal-Envelope Frontend with Learnable Per-Channel Energy Normalization for Whisper-Based Children's ASR

Temporal envelopes carry cues critical to speech intelligibility, yet ASR frontends based on log-mel spectrograms do not explicitly model continuous sub-band envelope structure. This limitation is particularly acute for children's speech, where high acoustic variability demands robust feature representations. We propos...

Edem Ahadzi, Ruchi Pandey, Tomi H. Kinnunen · 0 citations
Open access Aug 2026

Spectro-temporal vs. spectral features to predict the lombard gain in Mandarin Chinese

Background In background noise, speakers adapt their speech production, giving rise to Lombard speech, which often improves speech intelligibility (SI). While intelligibility benefits of Lombard speech have been extensively studied in non-tonal languages, it remains unclear whether spectro-temporal cues, which are crit...

M. Scharf, A. Warzybok, L. Wong et al. · 0 citations
Open access 2026

Automatic Detection of Misarticulation in Low-Resource Language Children Using Kaldi-Based ASR and Machine Learning Approaches

The lack of specific standards for evaluating the significant variation in children’s speech and the limited availability of annotated speech samples make it hard to automatically identify misarticulation in children speaking low-resource Indian languages. While techniques for assessing pronunciation using Automatic Sp...

Anushri Ghadge, S. Mahajan, Atharva Deshpande · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.