Machine Learning Research on Depression Assessment: A Narrative Review of Physiological, Behavioral, and Multimodal Approaches
Abstract
Depression has become an important target of machine learning research because conventional assessment often depends on intensive clinician-patient interaction and may be influenced by subjectivity. This narrative review examines recent studies on depression assessment from three major perspectives: physiological signals, behavioral signals, and multimodal fusion. After outlining the conceptual background of depression detection and prediction—including common datasets, methodological workflows, and evaluat ion issues—it compares representative studies using electroencephalography, heart-rate-related signals, speech, facial behavior, text, and combined modalities, highlighting how different data modalities capture complementary aspects of depressive states. A cross these domains, both traditional machine learning and deep learning methods have shown promising performance, with reported within-dataset accuracies often exceeding 90% and AUCs above 0.9 in some EEG and ECG studies. However, performance frequently d rops sharply in cross-site or multi - cohort settings—falling to around 52% balanced accuracy after harmonization in one large multi-site MRI benchmark—and major challenges remain in sample size, label quality, external validation, interpretability, fairness, and clinical deployment. Multimodal strategies appear especially promising, though they also introduce greater complexity in data collection and integration. Overall, future progress will depend not only on improving predictive accuracy but also on build ing reliable, explainable, and clinically usable systems for real-world mental health care.