Skip to content
Open access

Domain-informed density extraction for robust mental health classification of long social media posts

Aug 2026 · Frontiers in Artificial Intelligence · Vol 9 · 0 citations · 37 references
Medicine

TL;DR

The findings demonstrate that a leakage-safe evaluation protocol is essential for producing credible results in mental health classification from long social media posts and show that domain-informed density extraction provides a robust and practical text representation, with its greatest benefits emerging when models are unable to process the entire post.

Abstract

Background Recent advances in machine learning and natural language processing (NLP) have enabled the early identification of mental disorders from social media content. Among such platforms, Reddit is characterized by linguistically rich, long-form narratives in which users describe their psychological symptoms, emotions, and personal experiences. However, these datasets present significant challenges for text classification because the posts are lengthy, voluminous, and class-imbalanced. In this study, we revisit the concept of a segmentation-based, expression-weighted representation that emphasizes clinically relevant language while suppressing irrelevant content. During this investigation, we identify a critical limitation in conventional preprocessing pipelines—specifically, the application of chunking and oversampling prior to the train–test split—which introduces data leakage and consequently inflates reported performance. To address this issue, we propose a leakage-safe evaluation protocol together with a domain-informed density extraction method that identifies clinically dense passages using a lexicon derived exclusively from the training data. Results The proposed method was evaluated using logistic regression, linear SVM, XGBoost, fastText, RNN, and TextCNN, together with a transformer-based baseline (MentalBERT), all under a leakage-safe evaluation protocol. After eliminating data leakage, the apparent advantage of naive weighted chunking disappeared, and the headline macro-F1 score decreased from approximately 0.96 to 0.67. In contrast, the proposed domain-informed density extraction method consistently outperformed fixed-window chunking in five of the six models and matched or exceeded the full-text baseline in several cases. Statistically significant improvements were observed for logistic regression, RNN, and fastText, despite using only a fraction of the input text. For the context-limited transformer, density-based selection significantly outperformed naive truncation (macro-F1 + 0.023, 95% CI [+0.006, +0.041], p = 0.008). Conclusion Our findings demonstrate that a leakage-safe evaluation protocol is essential for producing credible results in mental health classification from long social media posts. They further show that domain-informed density extraction provides a robust and practical text representation, with its greatest benefits emerging when models are unable to process the entire post. By selectively preserving clinically informative content while reducing input length, the proposed approach offers an effective and computationally efficient alternative to increasingly complex model architectures.

Read PDF

Similar papers

Review Aug 2026

A Review of Deep Learning-Based Text Classification Research

The exponential growth of textual data on social media and information networks poses a significant challenge to extracting valuable information. Text classification, a core task in Natural Language Processing (NLP), is essential for organizing and categorizing such data. Deep learning has emerged as an effective a...

Ran Jin, Ya Wang, Tianzi Wu et al. · 0 citations
Open access Sep 2026

Modeling Co-Existing Mental Health Risks in Social Media via Multi-Label Learning and LLM

Mental health risk detection from user-generated social media text has become increasingly important as psychiatric conditions such as depression, anxiety, and suicidal ideation continue to rise. However, most existing computational studies operationalize the problem using single-label datasets, implicitly assuming mut...

Z. Özer, Emre Kenger, Görkem Gözükara et al. · 0 citations
Open access 2026

QBERT-LSTM: Quantum Intelligence-Based Mental Health Sentiment Analysis Using Web Scraping

A hybrid framework for sentiment classification from text, termed QBERT-LSTM, which integrates quantum-enhanced bidirectional encoder representations from transformers (QBERT) with long short-term memory (LSTM) networks, which excel at capturing global context and enhances sequential patterns and temporal features.

Najnin Sultana Shirin, Md. Aminul Islam, Maria Akter Abin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.