Skip to content
Open access

Modular speech-to-speech conversational system for Nigerian-accented English in healthcare

Sep 2026 · Nature Journal of Emerging Sciences Technologies and Innovations · 0 citations

Abstract

The performance of conversational AI systems tends to significantly decrease when used with accents which are not part of their initial training, which usually consists of the most popular English accents. This study proposes and evaluates a cascaded speech-to-speech framework tailored to Nigerian-accented English, in order to improve the accuracy and naturalness of such systems for Nigerian users. The system integrates a fine-tuned automatic speech recognition backbone based on Whisper, a natural language processor termed N-ATLAS, and a text-to-speech neural speech synthesis component implemented using Tacotron. Sequential domain adaptation was performed on OpenAI’s Whisper ASR using Common Voice (CV) and the African Accented Speech Dataset (AASD), each comprising approximately 3,400 utterances. The pretrained Whisper baseline achieved a Word Error Rate (WER) of 0.61 on AASD. Fine-tuning on CV reduced WER to 0.44, while subsequent adaptation on AASD further reduced WER to 0.28. In subjective end-to-end pilot evaluation, it was found that ASR accuracy and synthesis quality strongly correlate, as the intelligibility of the model improved substantially when WER fell below 0.30. The study contributes empirical evidence supporting sequential domain adaptation, hybrid symbolic-neural architectures, and modular cascaded speech-to-speech modelling for underrepresented accent varieties. The results highlight a scalable pathway toward more inclusive and adaptable speech technologies in low-resource contexts.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.