Skip to content
Review Open access

Using large language models to create lexicons for interpretable text models with high content validity: the Suicide Risk Lexicon

Sep 2026 · Journal of Psychopathology and Clinical Science · 0 citations
Medicine

TL;DR

This work demonstrates that LLMs –despite being black-boxes– can counterintuitively create interpretable models by generating lexicons, when this is preferred, and highlights the broader application of lexicons beyond measurement.

Abstract

Researchers often want to measure a variety of constructs such as anxiety, discrimination, or loneliness in text data from surveys, interviews, social media, and electronic health records. Using large language models (LLMs) –while optimal for text classification– remain infeasible for some researchers due to concerns around computational expertise, cost, privacy, and compute requirements. Therefore, some researchers prefer to use lightweight models for large datasets or interpretable models to avoid mistakes in high-stakes scenarios such as suicide risk detection. Lexicons offer simple baselines to LLMs by searching for relevant phrases –and can be used together with LLMs to guarantee capturing specific keywords in a deterministic way. However, building new lexicons is resource intensive. In this study, we found that GPT-4 turbo was able to automatically create a lexicon for 49 known risk factors for suicidal thoughts and behaviors, which we release as the Suicide Risk Lexicon. Generating a lexicon with LLMs quickly measures most constructs relevant for suicide risk detection, resulting in high content validity. This lexicon was able to accurately predict risk in crisis counseling conversations. After validating the lexicon with clinician ratings, the lexicon modestly outperformed the LIWC lexicon –which has low content validity for mental illness– and performed similarly to some black-box deep learning models. Due to using an interpretable approach with high content validity, we discovered that active suicidal ideation and direct self-injury were stronger indicators of imminent risk than passive suicidal ideation and depressed mood in this ecological setting (i.e., naturalistic counseling conversations as the crisis is happening, as opposed to assessments or interviews about retrospective symptoms). To simplify creating new lexicons for other research domains, we introduce a Python package, construct-tracker, that works with a variety of LLMs. In sum, while we recommend using LLMs for text classification, they remain out of reach for many researchers. Our work demonstrates that LLMs –despite being black-boxes– can counterintuitively create interpretable models by generating lexicons, when this is preferred. However, more studies are needed to test how these methods generalize to other datasets and psychological domains. Furthermore, we highlight the broader application of lexicons beyond measurement, including their use in benchmarking LLM performance.

Read PDF

Similar papers

Open access Aug 2026

The use of large language models in automated depression detection.

Current performance estimates of LLMs with respect to depression screening are most likely optimistic, but when restricted to smaller models that could be locally deployed (for privacy protection) in a clinical setting, LLMs do not detect depression with sufficient accuracy, sensitivity, or specificity to be used in a...

Sing-Hui Ling, W. Chorney · 0 citations
#natural language process... Preprint Aug 2026

Reading Anxiety or Reading the Label? Comparing Fine-Tuned and Frontier Models for Anxiety Detection on Social Media

Anxiety is among the most common mental health conditions, and people often write about it online well before seeking clinical help. Practitioners building detection tools face a concrete choice: call a frontier commercial model, fine-tune a smaller model in-house, or deploy a conventional classifier. We compare six co...

Cris Huynh, Arlene Pham · 0 citations
Aug 2026

Quantifying Social Biases in Language Model Classifiers is Domain-Dependent

This work investigates whether large language models (LLMs) can automatically adapt template-based bias datasets to specific domains using zero-shot prompting and shows that domain-adapted templates capture real-world bias patterns more faithfully than standard templates.

Tamara Quiroga, Felipe Bravo-Marquez, Valentin Barrière · 0 citations
#artificial intelligence Preprint Sep 2026

Bag of Tricks or Bag of Myths? Reducing Modeling Complexity with Task Knowledge in Explainable Suicide Risk Assessment

Assessing suicide risk from social media text is a small-data, high-stakes setting requiring not only severity prediction but also supporting evidence and clinically relevant risk and protective factors. Yet common NLP techniques, including model scaling, synthetic data, loss reweighting, ensembling, and threshold tuni...

Shlok Shelat, Shrey Salvi, Souvik Roy et al. · 0 citations
Preprint Aug 2026

Natural Language Processing Psychometrics

Results show the promise and limits of synthetic data: LLM personas can expose model biases, recover patterns consistent with clinical rumination, and support psychometric prediction from human text without a matched questionnaire, but cannot substitute for human validation.

Edoardo Sebastiano De Duro, Emma Franchino, M. Stella · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.