A decorrelated, clinician-aligned symptom signal readable directly from internal activations is revealed, offering a mechanistic foundation for interpretable depression-assessment tools.
Abstract
Patients with depression present with diverse symptom profiles, yet clinical practice routinely reduces this variation to a single severity score. Large language models (LLMs) can potentially capture various symptoms and their severity from patient speech. However, how depressive symptoms are represented inside LLMs remains poorly understood, limiting clinical trust. To examine whether internal model activations match clinician judgment, we analyzed the residual stream of Gemma-3-27B-PT using mechanistic interpretability techniques. Recording activations across symptom descriptions drawn from validated clinical instruments, we found that symptom groups geometrically separated the most at layer 21 across multiple distance metrics. Using Semantic Projection, we then projected held-out naturalistic text onto Symptom Vectors constructed from these instruments. The resulting per-symptom coefficients preserved clinician-annotated rank ordering across mood, somatic, and suicidality axes. Furthermore, a single depression vector in Layer 21 separates held-out depressive from non-depressive text (AUC = 0.789), which can be used as an emotional valence gate that restricts symptom projection to depressive speech. These results reveal a decorrelated, clinician-aligned symptom signal readable directly from internal activations, offering a mechanistic foundation for interpretable depression-assessment tools.
BACKGROUND
Remotely captured spoken language could provide objective, regular indicators of depression symptom severity. However, research to date has largely used non-clinical, cross-sectional written language and complex machine learning (ML) approaches with limited interpretability.
METHODS
We used linear mixed-ef...
A. Tokareva, J. Dineley, Z. Firth et al.· Journal of Affective Disorde...· 0 citations
Large Language Models (LLMs) are proposed as tools for high-throughput, deep phenotyping of psychiatric disorders. Applied to electronic health records, LLMs could in principle extract patient symptoms, outcome trajectories, risk factors, and treatment history at scale and these, when combined with increasingly availab...
S. Lock, J. Boisson, L. M. Evans et al.· medRxiv· 0 citations
This framework successfully decodes cognitive coping strategies independent of depressive mood, providing potentially actionable targets for personalized psychiatric interventions and a promising, interpretable digital biomarker for psychological resilience.
Shu-Yao Wang, Xue-Quan Zhu, Nan-Xi Li et al.· BMC Psychiatry· 0 citations
Abstract Background and Hypothesis Patients who are at clinical high risk (CHR) for schizophrenia need close monitoring of their symptoms to inform appropriate treatments. The Brief Psychiatric Rating Scale (BPRS) is a validated, commonly used research tool for measuring symptoms in patients with schizophrenia and othe...
Andrew X. Chen, G. Horga, Sean Escola· Schizophrenia bulletin· 1 citation
Using speech as objective markers for major depressive disorder (MDD) has shown promise, yet their generalizability across clinical settings remains largely unvalidated. This study aimed to validate previously identified speech markers of depressive symptoms in an independent clinical cohort, thereby assessing their re...
F. Menne, Felix Dörr, J. Tröger et al.· Annals of General Psychiatry· 0 citations
The usefulness of automated speech and language markers to monitor or predict psychotic symptoms depends on their ability to detect changes in mental state. To date, research linking psychosis and Natural Language Processing (NLP) has been conducted almost exclusively using cross-sectional experimental designs, lim...
S. Just, Shrankhla Pandey, D. Stein et al.· Translational Psychiatry· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 29, 2026
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.