Results show the promise and limits of synthetic data: LLM personas can expose model biases, recover patterns consistent with clinical rumination, and support psychometric prediction from human text without a matched questionnaire, but cannot substitute for human validation.
Abstract
Natural Language Processing (NLP) models predicting mental health outcomes rarely specify what they measure: contextual knowledge, emotional content, or syntactic structure. NLP Psychometrics treats psychological prediction from text as a psychometric problem, linking scores to interpretable linguistic evidence and testing beyond the training text format. Nine LLMs, conditioned on controlled personas (cognitive digital shadows), completed psychometric questionnaires with textual explanations per item. We extracted emotional profiles and syntactic-semantic structure via textual forma mentis networks, combined with personality and sociodemographic variables in ablated random forest (RF) regressors, using SHAP to identify which features drove performance and in which direction. Full RF models explained up to 70.8% of variance in life satisfaction (SWLS), 55.7% in depression (PHQ-9), and, for DASS-21, 68.5% depression, 76.0% anxiety, 72.4% stress. Sociodemographics alone explained no meaningful variance in depression, anxiety, or stress, but did so for life satisfaction, where emotion features and income were the strongest predictors; neuroticism and network topology instead dominated depression and anxiety, reversing direction between them. Without retraining, RF models separated diaries from low- and high-score personas ($r$ up to 0.91) and, using only network/emotion features, classified clinical from control participants in real transcripts with up to 68% accuracy. These results show the promise and limits of synthetic data: LLM personas can expose model biases, recover patterns consistent with clinical rumination, and support psychometric prediction from human text without a matched questionnaire, but cannot substitute for human validation. NLP Psychometrics makes these distinctions explicit, measurable, and testable through interpretable AI and network/emotional features.
BACKGROUND
Remotely captured spoken language could provide objective, regular indicators of depression symptom severity. However, research to date has largely used non-clinical, cross-sectional written language and complex machine learning (ML) approaches with limited interpretability.
METHODS
We used linear mixed-ef...
A. Tokareva, J. Dineley, Z. Firth et al.· Journal of Affective Disorde...· 0 citations
Social media provides a valuable source for early mental health detection using natural language processing. Although BERT-based models can identify subtle psychological signals in text, most studies focus on English data, while the effect of automated translation on Indonesian depression detection remains underexplore...
Muhamad Sandi Alfarizi, Adinda Mutiara Suci, A. A. S. Gunawan et al.· 2026 International Conferenc...· 0 citations
The usefulness of automated speech and language markers to monitor or predict psychotic symptoms depends on their ability to detect changes in mental state. To date, research linking psychosis and Natural Language Processing (NLP) has been conducted almost exclusively using cross-sectional experimental designs, lim...
S. Just, Shrankhla Pandey, D. Stein et al.· Translational Psychiatry· 0 citations
This work demonstrates that LLMs –despite being black-boxes– can counterintuitively create interpretable models by generating lexicons, when this is preferred, and highlights the broader application of lexicons beyond measurement.
Daniel M. Low, Osiris Rankin, Daniel D. L. Coppersmith et al.· Journal of Psychopathology a...· 0 citations
Natural language processing (NLP) has progressed from narrow, task-specific statistical models to large language models (LLMs) capable of generating fluent, contextually coherent text across virtually unlimited domains, a transition that has reshaped three previously distinct application areas simultaneously: education...
R. Rakesh· Natural Resources for Human...· 0 citations
A decorrelated, clinician-aligned symptom signal readable directly from internal activations is revealed, offering a mechanistic foundation for interpretable depression-assessment tools.
Fang-Ying Zhu, Ajay K. Subramanian, A. Constant et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.