Skip to content
Open access

Early Stage Depression Recognition Using Machine Learning and Natural Language Processing: A Hybrid PHQ-9 and Linguistic Feature-Based Screening Framework

Aug 2026 · International Journal of Innovative Science and Research Technology · 0 citations · 6 references

Abstract

Depression is a widespread and frequently under-diagnosed mental health condition, and delays in identification are associated with poorer long-term outcomes. This paper presents a hybrid screening framework that combines a validated questionnaire to the Patient Health Questionnaire-9 (PHQ-9), with linguistic feature extraction from short free-text journal entries to generate a composite depression-risk score. Unlike approaches that rely on a single data modality, the proposed system fuses structured clinical scoring with textual indicators associated with depressive symptomatology in prior computational-linguistics research, such as elevated first-person pronoun usage, negation patterns, and lexical markers of sadness, anhedonia, and worthlessness. A prototype web application implementing this framework was developed, comprising a screening interface, a relational database schema for longitudinal storage, and an analytics dashboard for aggregate monitoring. This paper describes the system architecture, the feature-fusion methodology, the underlying database design, and a discussion of how the prototype's rule-based linguistic module can be replaced with a trained supervised classifier (e.g., Support Vector Machine, Random Forest, or a fine-tuned transformer model) in future work. The framework is intended as a low-cost, scalable screening aid to support — not replace — clinical [1] evaluation.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.