Skip to content
Preprint

Natural Language Processing Psychometrics

Aug 2026 · 0 citations · 88 references
Computer Science

TL;DR

Results show the promise and limits of synthetic data: LLM personas can expose model biases, recover patterns consistent with clinical rumination, and support psychometric prediction from human text without a matched questionnaire, but cannot substitute for human validation.

Abstract

Natural Language Processing (NLP) models predicting mental health outcomes rarely specify what they measure: contextual knowledge, emotional content, or syntactic structure. NLP Psychometrics treats psychological prediction from text as a psychometric problem, linking scores to interpretable linguistic evidence and testing beyond the training text format. Nine LLMs, conditioned on controlled personas (cognitive digital shadows), completed psychometric questionnaires with textual explanations per item. We extracted emotional profiles and syntactic-semantic structure via textual forma mentis networks, combined with personality and sociodemographic variables in ablated random forest (RF) regressors, using SHAP to identify which features drove performance and in which direction. Full RF models explained up to 70.8% of variance in life satisfaction (SWLS), 55.7% in depression (PHQ-9), and, for DASS-21, 68.5% depression, 76.0% anxiety, 72.4% stress. Sociodemographics alone explained no meaningful variance in depression, anxiety, or stress, but did so for life satisfaction, where emotion features and income were the strongest predictors; neuroticism and network topology instead dominated depression and anxiety, reversing direction between them. Without retraining, RF models separated diaries from low- and high-score personas ($r$ up to 0.91) and, using only network/emotion features, classified clinical from control participants in real transcripts with up to 68% accuracy. These results show the promise and limits of synthetic data: LLM personas can expose model biases, recover patterns consistent with clinical rumination, and support psychometric prediction from human text without a matched questionnaire, but cannot substitute for human validation. NLP Psychometrics makes these distinctions explicit, measurable, and testable through interpretable AI and network/emotional features.

View source

Similar papers

Open access Sep 2026

Multilingual lexical feature analysis of spoken language for predicting major depression symptom severity.

BACKGROUND Remotely captured spoken language could provide objective, regular indicators of depression symptom severity. However, research to date has largely used non-clinical, cross-sectional written language and complex machine learning (ML) approaches with limited interpretability. METHODS We used linear mixed-ef...

A. Tokareva, J. Dineley, Z. Firth et al. · 0 citations
Conference Aug 2026

Depression Datasets: Does Language Matter? A Comparative Study of English BERT and IndoBERT on Translated Depression Datasets

Social media provides a valuable source for early mental health detection using natural language processing. Although BERT-based models can identify subtle psychological signals in text, most studies focus on English data, while the effect of automated translation on Indonesian depression detection remains underexplore...

Muhamad Sandi Alfarizi, Adinda Mutiara Suci, A. A. S. Gunawan et al. · 0 citations
Open access Aug 2026

Changes in speech reflect changes in psychotic symptom severity: a longitudinal natural language processing analysis

The usefulness of automated speech and language markers to monitor or predict psychotic symptoms depends on their ability to detect changes in mental state. To date, research linking psychosis and Natural Language Processing (NLP) has been conducted almost exclusively using cross-sectional experimental designs, lim...

S. Just, Shrankhla Pandey, D. Stein et al. · 0 citations
Review Open access Sep 2026

Using large language models to create lexicons for interpretable text models with high content validity: the Suicide Risk Lexicon

This work demonstrates that LLMs –despite being black-boxes– can counterintuitively create interpretable models by generating lexicons, when this is preferred, and highlights the broader application of lexicons beyond measurement.

Daniel M. Low, Osiris Rankin, Daniel D. L. Coppersmith et al. · 0 citations
Review Sep 2026

Natural Language Processing and Large Language Models: AI for Education, Communication, and Digital Content Analysis

Natural language processing (NLP) has progressed from narrow, task-specific statistical models to large language models (LLMs) capable of generating fluent, contextually coherent text across virtually unlimited domains, a transition that has reshaped three previously distinct application areas simultaneously: education...

R. Rakesh · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.