Skip to content

Multimodal Fusion of Physiological Signals and Behavioral Data for Mental Health Assessment: A Deep Learning Approach Based on Hierarchical CNN

Aug 2026 · Journal of Mechanics in Medicine and Biology · 0 citations

Abstract

The accurate and objective recognition of mental health conditions, such as anxiety and depression, remains a significant challenge in biomedical engineering and clinical practice, where traditional assessment methods often suffer from subjectivity, inefficiency, and a lack of quantitative biomarkers. In this study, a hierarchical multimodal convolutional neural network (CNN) is proposed. By integrating heterogeneous biomedical data sources from different public datasets, including texts, facial images, and physiological signals, a multimodal learning framework is constructed for mental health status recognition. Special convolution and pooling structures are designed for different modalities, fine-grained spatial and temporal features are extracted from corresponding data sources, and the complementary information between different modal combinations is learned through a dynamic fusion mechanism. Experimental evaluations are conducted on the publicly available Distress Analysis Interview Corpus - Wizard of Oz (DAIC-WOZ) dataset and the Affect, Mood, and Personality Database (AMIGOS). On the DAIC-WOZ dataset, the proposed model achieved an accuracy of 0.818, a recall of 0.801, an F1 score of 0.809, an area under the curve (AUC) of 0.868, and a precision-recall (PR)-AUC of 0.842. Compared with the best-performing baseline (Multimodal Transformer), it increases by 6.5%, 7.2%, 6.9%, 6.1%, and 8.5%, respectively. The model's effectiveness in multimodal feature extraction and fusion is verified on the AMIGOS dataset. The proposed model is evaluated in the emotion recognition task. The accuracy under short and long video conditions reaches 0.791 and 0.801, respectively, and the corresponding AUC values are 0.837 and 0.849. Ablation experiments confirm that multimodal feature fusion significantly enhances recognition performance compared to single-modality approaches. In contrast, robustness tests demonstrate stable performance across varying hyperparameter configurations. These results indicate that the proposed hierarchical dynamic multimodal CNN framework can effectively integrate text, visual, and physiological information based on the modal features of different datasets; thus, it can provide a repeatable, objective, and clinically feasible technical approach for intelligent mental health assessment and personalized intervention.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.