Skip to content
Open access

Dual-modality modeling for depression detection using speech signals

Sep 2026 · Scientific Reports · 0 citations

Abstract

Depression is a leading contributor to the global burden of disease and a significant barrier to both personal well-being and societal development. Yet, it remains underdiagnosed, primarily due to social stigma and limited access to clinical resources. This study proposes a dual-modality framework for automated depression detection using speech signals, capturing both what individuals say (text) and how they say it (vocal cues) from speech recordings. The method integrates contextual embeddings and acoustic features extracted from clinical interview recordings. Transcriptions are generated using the Whisper automatic speech recognizer, and text embeddings are derived from a fine-tuned BERT model, while prosodic and spectral features serve as vocal cues. To improve classification sensitivity under class imbalance, random oversampling is applied during fine-tuning. Both feature-level and decision-level fusion strategies are evaluated using fully connected and Long Short-Term Memory (LSTM) networks. On the E-DAIC dataset, the decision-level OR fusion strategy achieved 96.68% accuracy, 90.41% precision, 95.17% recall, and a 92.73% F1-score. Evaluation on the Chinese MODMA dataset assesses the applicability of the framework to an additional language-specific dataset after dataset-specific training, while indicating that modality contribution may vary across datasets. The reported results are segment-level methodological findings and should not be interpreted as participant-independent clinical performance. The speech-only design avoids visual-data collection but does not eliminate the privacy risks associated with identifiable speech recordings and transcripts.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.