Machine Learning based Depression Detection with NLP Applicationon Twitter Text Data
Abstract
Depression detection through social media analysis has gained increasing attention due to the accessibility of large-scale user-generated content. This study proposes a hybrid sentiment analysis framework that integrates Latent Dirichlet Allocation (LDA), Valence Aware Dictionary and Sentiment Reasoner (VADER), and Term Frequency–Inverse Document Frequency (TF-IDF) features to capture thematic, emotional, and statistical representations of text. The dataset, sourced from Twitter, was preprocessed with advanced filtering techniques to remove noise and irrelevant tokens, producing a robust feature space. ML classifiers are evaluated like Support Vector Machine (SVM), K-Nearest Neighbor (KNN), logistic regression, and random forest, and SVM achieved the best accuracy (95.1%), precision (98.8%), recall (89.4%), and F1-score (93.9%). The results show that the combination of probabilistic, lexicon-based and statistical approaches improves classification results over the individual approaches and provides an interpretable, scalable framework for computational depression detection.