Skip to content
Open access

Diagnostics of speech disorders in psychological and speech therapy practice: crosslingual validation of the SlowFast model

2026 · Sovremennaya nauka i innovatsii · pp. 210-221 · 0 citations

TL;DR

The results confirm the high cross-lingual transferability of the SlowFast architecture and open prospects for creating multilingual automated speech disorder diagnostic systems new language.

Abstract

In the authors’ previous study, the SlowFast two-stream architecture demonstrated high efficiency in diagnosing dyslalia on Russian-language data (98.0% accuracy on real clinical recordings). However, the question of the model's cross-lingual robustness remained open. This paper evaluates the ability of a SlowFast model trained on Russian to correctly classify speech disorders in Polish-speaking children. The open PAVSig dataset (N=201 children, 66,781 audio segments with double expert diagnosis of sigmatism) was used as the target corpus. In zero-shot evaluation mode (without any retraining on Polish data), the without the need to collect large labeled corpora for each model achieved 87.3% accuracy, 86.8% precision, 88.1% recall, 87.4% F1-score, and 0.92 AUC-ROC. Minimal adaptation (finetuning on 10% of the data) increased accuracy to 91.8%, and on 25% of the data – to 94.2%. Error analysis revealed that the main difficulties are associated with acoustic differences in Polish sibilant sounds (sz, cz, ż), accounting for 58% of all errors in zero-shot mode, as well as background noise (27%) and short segment duration (15%). After fine-tuning on 25% of the data, the proportion of errors due to cross-lingual acoustic differences decreased to 22%. The results confirm the high cross-lingual transferability of the SlowFast architecture and open prospects for creating multilingual automated speech disorder diagnostic systems new language.

Read PDF

Similar papers

#machine learning Preprint Sep 2026

Deep Learning Techniques for Phoneme Recognition in Italian Children's Speech

Speech therapists often face difficulties diagnosing impairments due to the lack of efficient tools for transcribing speech into the International Phonetic Alphabet (IPA). This work addresses this challenge with Broca, a Conformer-based deep learning system pretrained on 8 days of adult speech and fine-tuned on a 165-m...

N. Barbaro, Cristina Gena, F. Petriglia et al. · 0 citations
Open access Aug 2026

Model Development and Validation for Repetition Severity Assessment in Stuttered Speech Using Clinical Speech Datasets

A computational model that grades repetition severity from clinical speech recordings, evaluated on 480 audio samples from 60 adult speakers with persistent developmental stuttering, supports its use as a clinical decision-support tool.

J. N. Pooja, H. Y. Vani, R. P et al. · 0 citations
Open access Aug 2026

The Predictive Role of Specific Audiometric Frequencies in Speech-in-Noise Perception: A Machine Learning Approach

Purpose: Speech comprehension performance in noisy environments is a complex process that cannot be fully predicted by standard audiometric assessments. The aim of this study is to use machine learning (ML) algorithms to predict individuals’ difficulties in understanding speech in noise based on clinical data.Methods:...

Özgenur Gavgalı · 0 citations
Open access Sep 2026

Acoustic Features of Sustained Phonation for Schizophrenia Classification: A Feasibility Study

Voice and speech are increasingly studied as indicators of mental health, but the acoustic features of sustained phonation in schizophrenia remain underexplored. This feasibility study examined whether a 500 ms sustained vowel /a/ contains meaningful information for distinguishing patients with schizophrenia from healt...

L. Jelić, K. Jambrošić, Vinko Lešić et al. · 0 citations
Preprint Aug 2026

Motor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease Detection

A layer-wise analysis of nine SSL speech backbones using a low-capacity logistic regression probe reveals that the transferred discriminative signal lacks pathological specificity, highlighting critical limitations that must be addressed before speech-based pathology recognition models can be reliably deployed in clini...

Serli Kopar, Sam Gijsen, Abner Hernandez et al. · 0 citations
Open access Aug 2026

Automatic analysis of speech representations to assess psychological distress

Background Current mental health diagnostic methods are limited by subjective clinical interpretation. Automatic speech analysis is a promising technology for objective assessment. Objective To evaluate and compare different speech-based representations (acoustic, phonetic, and time-frequency) and deep learning-based e...

Sara Fernández-Velasco, Jose Moreno-Mesa, D. Escobar-Grisales et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.