Jul 2026· JUTI: Jurnal Ilmiah Teknologi Informasi· pp. 218-235· 0 citations· 36 references
TL;DR
This study explores the linguistic markers of depression and anxiety in Indonesian social media text through an integrated model of machine learning classification and topic modeling to suggest that combining classification with topic modeling offers a practical foundation for developing early detection tools and Indonesian-language NLP resources for mental health discourse analysis.
Abstract
The increasing prevalence of mental health disorders such as depression and anxiety calls for effective approaches to analyze psychological expressions in textual data. This study explores the linguistic markers of depression and anxiety in Indonesian social media text through an integrated model of machine learning classification and topic modeling. In contrast to earlier work primarily centered on classification performance, this work emphasizes interpretability through comparative machine learning analysis and LDA-based thematic analysis. Classification determines the expressed condition, LDA determines thematic structures that account for distinguishing patterns beyond accuracy metrics alone. The dataset consisted of 17,096 records collected from Facebook groups, reduced to 7,199 instances after removing neutral labels. TF-IDF-based feature extraction was applied using unigram and bigram representations with a maximum of 5,000 features. A comparative analysis was conducted using six classification algorithms: K-Nearest Neighbor (KNN), Support Vector Machine (SVM), Random Forest, Decision Tree, XGBoost, and Naïve Bayes, evaluated under three train-test split scenarios (70:30, 80:20, and 90:10). SMOTE was applied to the training data to address class imbalance. SVM achieved the best performance with an F1-score of 0.903 under an 80:20 split, followed by Naïve Bayes and XGBoost, while KNN performed lowest consistently. LDA topic modeling revealed that depression-related texts were dominated by internal emotional expression, social isolation, and suicidal ideation, whereas anxiety-related texts were characterized by sudden fear, somatic physical symptoms, and social anxiety. These findings suggest that combining classification with topic modeling offers a practical foundation for developing early detection tools and Indonesian-language NLP resources for mental health discourse analysis.
Depression is a rising mental health concern among college students, often manifesting through linguistic and behavioural cues on social media platforms such as Facebook. This study aims to analyse social media comments using text mining techniques to detect potential signs of depression. The research applies a structu...
Mental disorders such as depression constitute the largest share of concerns in mental well-being across the globe, having a severe influence on an individual's emotional, social and professional lives. With the increased application of social networks in the modern age, opportunities have opened up to detect the signs...
Priyanka Srivastava, Mohammad Suaib· Natural Resources for Human...· 0 citations
It is demonstrated that machine learning algorithms can serve as effective tools for classifying Facebook text into varying levels of depression severity, but challenges such as class imbalance, limited data availability, and discrepancies between self-assessment results and online behavior indicate that further resear...
Keito R. Yoneyama· Chulalongkorn Medical Journa...· 0 citations
Depression detection through social media analysis has gained increasing attention due to the accessibility of large-scale user-generated content. This study proposes a hybrid sentiment analysis framework that integrates Latent Dirichlet Allocation (LDA), Valence Aware Dictionary and Sentiment Reasoner (VADER), and Ter...
Priyanka Srivastava, Mohammad Suaib· Natural Resources for Human...· 0 citations
Social media and digital platforms have changed the landscape of mental health discourses, making platforms such as Twitter key for individuals to discuss their mental health and find solace. Although these platforms provide important linguistic insights for different mental health disorders, most current computational...
Misha M, Hamza Muneer, R. Tehseen et al.· International Journal of Inn...· 0 citations
The research stages include a preprocessing process consisting of cleaning, case folding, tokenizing, normalization, stopword removal, and stemming. Furthermore, the TF-IDF method is used to extract text features so that the data can be represented in numerical form. The dataset was then divided into training data and...
Olinda Nathaniel Mendrofa, Deva Angriani, E. Ompusunggu· Formosa Journal of Computer...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.