Skip to content
Conference

From Lived Experience to Digital Biomarkers: Methods for Scaling Narrative Data with Large Language Models in Mental Health

Jul 2026 · International Conference on Digital Health · pp. 450-452 · 0 citations · 8 references

Abstract

Lived experience narratives provide a rich account of how individuals interpret and organize their mental health, capturing dimensions of meaning and context that are often missed by structured assessments and traditional language-based features. Despite their importance, the unstructured nature of these data has limited their systematic use. Recent advances in large language models (LLMs) enable scalable analysis of narrative data, allowing for the extraction of thematic and structural features across large datasets. These approaches position lived experience as a promising digital biomarker, with the potential to capture early, ecologically valid signals and complement existing clinical measures. However, challenges related to validation, interpretability, bias, and data governance remain. We outline emerging methodological frameworks and discuss how LLMbased approaches can support more scalable, longitudinal, and person-centered models of mental health.

View source

Similar papers

Review Open access Aug 2026

Generative Large Language Models in Mental Health Care Settings: Systematic Review and Meta-Analysis

Abstract Background General-purpose large language models (LLMs) are increasingly being tested in mental health care, where language is central to assessment, diagnosis, risk evaluation, therapeutic interaction, monitoring, and patient education. However, their clinical usefulness, safety, and readiness for implementation remain uncertain. Existing reviews have largely been descriptive or scoping in nature, and broad health care reviews have not examined in detail the distinctive risks and applications of LLMs in mental health care. Objective We aim to systematically review empirical evidence on the use of general-purpose LLMs in mental health care; characterize the clinical tasks, study designs, models, outcomes, and methodological quality of the evidence; and synthesize findings across clinically meaningful task domains, including quantitative synthesis where sufficiently comparable studies were available. Methods We followed PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) and searched PubMed, Embase, ACM Digital Library, IEEE Xplore, and Google Scholar from November 2022 to March 2026. Eligible studies reported quantitative data evaluating general-purpose LLMs for direct mental health care tasks. Methodological quality was assessed using the Mixed Methods Appraisal Tool and certainty of evidence using GRADE (Grading of Recommendations, Assessment, Development, and Evaluation) domains. Findings were narratively synthesized, and random-effects meta-analyses were conducted on studies that tested LLMs on screening and diagnosis by ChatGPT-4 (OpenAI), ChatGPT-3.5 (OpenAI), and GPT-3 (OpenAI) models for mental health outcomes that reported on specificity and sensitivity. Hartung-Knapp adjustments were applied, and prediction intervals (PIs) estimated. Results We included 66 studies, comprising 37 vignette or simulation studies, 22 retrospective studies, and 7 prospective studies. Applications included screening and diagnosis (n=29), clinical decision support (n=14), treatment support (n=10), documentation and monitoring (n=6), patient education (n=4), and patient engagement (n=3). In screening and diagnosis, meta-analysis of 8 studies found that GPT-4 had the higher pooled sensitivity (0.83, 95% CI 0.38‐0.97; 95% PI 0.02‐1.00) than GPT-3.5 (0.70, 95% CI 0.13‐0.97; 95% PI 0.00‐1.00) and GPT-3 (0.61, 95% CI 0.33‐0.82; 95% PI 0.10‐0.96). However, GPT-4 specificity was lower at 0.77 (95% CI 0.52‐0.91; 95% PI 0.10‐0.99). Narrative synthesis suggested that LLMs performed most consistently in structured and linguistically explicit tasks. Certainty of evidence was generally low across domains, although documentation and monitoring reached moderate certainty. Major limitations included indirectness from vignette-based designs, uncertain representativeness, inconsistent outcome reporting, and sparse prospective real-world evaluation. Conclusions General-purpose LLMs show promise for selected mental health care applications. However, current evidence remains too heterogeneous, indirect, and uncertain to support routine unsupervised use, particularly for diagnosis, risk assessment, crisis response, or therapeutic interaction. Broad accessibility should not be mistaken for clinical readiness. Future studies should move beyond simulations and retrospective evaluations toward prospective, real-world research assessing safety, reliability, equity, acceptability, clinical outcomes, and implementation in diverse mental health care settings.

Janni Leung, Benjamin Johnson, Kelsey McRae et al. · 1 citation
Open access Jul 2026

A bibliometric analysis of large language models in mental health research

The analysis revealed a rapid acceleration in scholarly output, with a compound annual growth rate of 140%, driven by advancements in models such as GPT-3 and GPT-4, alongside strategic funding and industry initiatives.

Mohammad Ali Hussiny, T. Saidi, Minna Pikkarainen et al. · 1 citation
Conference Aug 2026

Demo: SNAQ: Quantifying Patient Narratives to Strengthen Longitudinal Assessment in Chronic Pain

Chronic pain assessment relies heavily on patients' descriptions of their symptoms, yet these narrative reports are difficult for clinicians to interpret consistently and are rarely incorporated into structured monitoring tools. The purpose of this research was to develop a computational system, SNAQ (System of Narrative Aspect Quantification), that translates patient narratives into quantitative indicators that can be tracked over time alongside standard symptom ratings. SNAQ uses natural language processing, specifically aspect-based sentiment analysis, to identify meaningful themes in patients' written descriptions and classify them using the World Health Organization's International Classification of Functioning (ICF). These sentiment-based measures are then combined with normalized 0–10 symptom scales to produce a single Wellness Index ranging from 0 to 100. We evaluated SNAQ using a synthetic dataset modeled on real fibromyalgia narratives and a six-month, 50-entry longitudinal case. SNAQ accurately identified functional themes and emotional tone in narratives, showing strong agreement with human reviewers (F1 = 0.855, κ = 0.673), and produced a stable index that reflected realistic patterns of symptom flare-ups and recovery. SNAQ is now being clinically validated on real patient data in an orthopedic practice, where its index is compared against an established PROM across multiple visits per patients to asses concurrent validity and response to clinical change. The study demonstrates that SNAQ functions as intended: it reliably extracts measurable signals from patient narratives and integrates them with symptom scales into a stable, clinically plausible Wellness Index, offering a transparent, interpretable tool for both patients and clinicians.

Ishaan Sinha, A. Sherafati · 0 citations
Review Open access Aug 2026

Application of Large Language Models in Chronic Disease Care: Mixed Methods Systematic Review and Thematic Synthesis

This review is the first to integrate Orem’s Self-Care Theory into a 3D evaluative framework for large language models in chronic care, thereby moving beyond fragmented, technology-centric assessments toward a structured, nursing-informed, and theory-driven synthesis.

Ling-Hui Zhang, P. Huai, Rui Xu et al. · 0 citations
Review Open access Aug 2026

Large Language Models for Mental Health Prediction: Scoping Review of Bias and Clinical Utility Documentation in 2019-2024

Abstract Background A growing body of literature leverages large language models (LLMs) to make mental health predictions. However, these models are prone to bias, and studies to validate their clinical utility are lacking. Objective This scoping review aims to uncover bias and clinical utility limitations stemming from the methodological design of LLM-based mental health predictive systems. In addition, it intends to document the level of self-reflection about bias and clinical challenges reported by authors in their own work. Methods This work follows the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines and was registered online. Eligible studies were original research articles in English published between 2019 and 2024, using LLMs to detect mental health conditions in nonsynthetic textual data. The search was conducted in 5 scientific databases (PubMed, Web of Science, IEEE Xplore, ACM Digital Library, and ACL Anthology) with queries associating keywords related to “Mental Health,” “Large Language Models,” and “Prediction.” We extracted both methodological information about the included studies and authors’ statements relevant to issues of bias and clinical utility. This extraction was based on a framework screening the entire pipeline of development of LLMs with applications in mental health: research design and selection, data collection, outcome definition, model development, and postdeployment considerations. Statistical description of the retrieved entities, as well as thematic coding, was performed for analysis. Results A total of 2472 articles were identified, of which 263 (10.6%) were assessed for eligibility, and 201 (8.1%) were included in the review. Included studies were mostly recent, indicating a growing interest in the use of LLMs for mental health predictions. Our analysis revealed that a majority of studies share similar methodological choices along their development pipeline: most of them focus on depressive disorders identified via processing user texts on social media, mainly with the use of nonspecialist LLMs derived from BERT (Bidirectional Encoder Representations from Transformers). Following previous works on these matters, we highlighted how these choices may hinder the clinical relevance and fairness of the envisioned systems. Similarly, we found that 164 (81.6%) studies mention themes related to bias and clinical utility; however, most of the discussion revolves around data-centered issues. Only 41 (20.4%) articles mention themes associated with at least 3 out of 5 pipeline steps, suggesting a limited appropriation of the notions of bias and clinical utility in such a sensitive context as mental health analysis. Conclusions Bias and clinical utility are lightly covered in the field of LLM-based mental health prediction research as of 2019‐2024. In-depth approaches involving interdisciplinary teams of clinicians and natural language processing specialists are needed to ensure technical soundness, clinical relevance, and fair outcomes for potential users.

Clémentine Bleuze, Karen Fort, Vincent P. Martin et al. · 0 citations
Review Open access Sep 2026

Adapting the narrative engagement scale for unstructured patient audio: a tool for advancing patient-centered research

Patient narratives often aim to foster empathy, sustain attention, and translate lived experiences into patient-centered research priorities; however, while the content of these stories has been widely examined, audience engagement, which can influence whether narratives achieve these aims, has received less attention. Currently, a standardized method to assess engagement with unstructured patient narratives is lacking. We aimed to validate an adapted narrative engagement scale and examine potential differences in engagement based on storyteller demographics and health topics. We analyzed audio narratives from patients and caregivers collected by [Platform Name Blinded for Review]. Participants ( n  = 646) shared unstructured health-related stories. We adapted the Busselle and Bilandzic narrative engagement scale, originally designed for visual media, to assess three dimensions: narrative understanding, attentional focus, and emotional engagement. Confirmatory factor analysis (CFA) was conducted on a subset of narratives ( n  = 174) to verify the factor structure. For the remaining narratives ( n  = 472), we calculated factor scores and performed Mann-Whitney U tests to compare engagement across demographic groups and health topics. Most participants were female, White, and college-educated. The adapted scale demonstrated preliminary evidence of acceptable psychometric performance. Overall, stories were rated as easy to understand and able to maintain audience focus. Demographic factors were not significantly associated with engagement scores. Notably, narratives of COVID-19 were significantly easier ( p <0.05) to understand and pay attention to, but less emotionally engaging than non-COVID-19 stories, whereas injury- or wound-related narratives elicited higher emotional engagement. Engagement scores remained consistent across other health topics. This study provides preliminary validation evidence for the adapted narrative engagement measure for evaluating audience engagement with unstructured patient audio narratives across diverse health concerns. Future efforts should focus on capturing narratives from underrepresented groups to strengthen generalizability. This adapted measure can help researchers systematically evaluate not only what patient stories communicate, but also how audience understand, attend to, and emotionally engage with those stories, supporting the meaningful use of patient narratives in patient-centered research.

Young-Ji Lee, K. Abebe, Kwonho Jeong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.