Machine Learning Prediction of Depression Risk from Behavioural and Self-Reported Data
Abstract
. The issue of depression has become a significant social concern in the world and there is need to use scalable methods to detect the risks at early onset. The paper explores how behavioural and demographic information predicts depression risk with the aid of several machine learning algorithms. Depression risk is identified using PHQ-9 screening tools, using NHANES survey data. The logistic regression, random forest, and gradient boosting are applied to assess the predictive performance of the behavioural-only and combined features. These findings demonstrate that behavioural variables by themselves have good predictive capacity as ROC-AUC is more than 0.80 in all models. Consistent but modest enhancements are given when the demographic variables are included. These results indicate that the measured lifestyle indicators can be used as informative variables to predict the presence of depression and outline the promise of machine learning strategies as scalable mental health screening of populations.