Skip to content
Conference

Depression Datasets: Does Language Matter? A Comparative Study of English BERT and IndoBERT on Translated Depression Datasets

Aug 2026 · 2026 International Conference on Smart Data, Intelligence, and Analytics (ICoSDIA) · pp. 1-6 · 0 citations · 30 references

Abstract

Social media provides a valuable source for early mental health detection using natural language processing. Although BERT-based models can identify subtle psychological signals in text, most studies focus on English data, while the effect of automated translation on Indonesian depression detection remains underexplored. This study compares multiclass detection of major depressive, bipolar, postpartum, psychotic, atypical, and no-depression categories using 14,983 text samples. It contrasts an English BERT model trained on original datasets with IndoBERT applied to datasets translated via the Google Translate API. Experimental findings indicate that English BERT attained an accuracy of $90.0 \%$ and a macroaveraged F1-Score of 0.91, but IndoBERT showed a significant performance deterioration, achieving merely 79.0% accuracy and a 0.79 F1-Score. A significant decline was noted in the psychotic group, where diagnostic recall plummeted from 0.87 in the original English dataset to 0.42 in the translated data. These findings experimentally suggest that language significantly influences mental health text classification. Translation artifacts were discovered to neutralize significant verbal cues and emotional keywords, consequently undermining diagnostic accuracy. This study highlights that original Indonesian datasets are crucial for highly dependable mental health monitoring, as the findings suggest that machine translation may introduce semantic shifts that affect the subtle linguistic context necessary in critical clinical applications.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.