Skip to content
Book Open access

Designing a Human-in-the-Loop LLM Chatbot for Interactive Depression Screening

Jul 2026 · Information Hiding · 0 citations · 24 references
Computer Science

TL;DR

The results suggest that the chatbot is a usable alternative for administering screening instruments and that process-level behavioral traces may add contextual information beyond final questionnaire responses.

Abstract

Depression remains a major global mental health challenge, and early screening still depends largely on self-report questionnaires, despite their vulnerability to response bias and limited insight into the response process. This study examined the feasibility of a human-in-the-loop, LLM-based chatbot for depression screening under real-time supervision by a licensed psychologist. Twenty-one participants completed the DASS-21 in two digital formats: a static form and a conversational chatbot interface. In the chatbot condition, LLM-generated follow-up questions were reviewed by the psychologist before delivery. Interaction data were collected in both conditions using a digital phenotyping screening tool. The results suggest that the chatbot is a usable alternative for administering screening instruments and that process-level behavioral traces may add contextual information beyond final questionnaire responses. These findings support the potential of clinician-supervised chatbot workflows for enriching depression screening, while highlighting the need for further validation of the clinical relevance of such behavioral signals.

Read PDF

Similar papers

Open access 2026

The DeepSeek chatbot as a tool for psychological support in processing significant life events

This article presents a pilot study investigating the effectiveness of DeepSeek chatbot as a tool for psychological support during significant life events. The relevance of this research is underscored by the increasing demand for accessible mental health resources. The study involved 64 participants, divided into two groups: an experimental group (n=39) and a control group (n=25). A quasi-experimental pre-test/post-test design was implemented over a four-week intervention period. Participants in the experimental group engaged in interactions with DeepSeek chatbot, while those in the control group received psychoeducational materials. To assess the intervention’s effectiveness, standardized instruments were used: the Social Readjustment Rating Scale (SRRS), the Depression, Anxiety, and Stress Scale (DASS-21), the Satisfaction with Life Scale (SWLS), and the UCLA Loneliness Scale. A content analysis of significant life events experienced by the participants in the past year revealed a similar structure of dominant stressors across both groups. Quantitative data analysis employed descriptive statistics, the Shapiro-Wilk test, Levene’s test, and Student’s test. The subjective experience of DeepSeek chatbot users was examined through a qualitative analysis of structured reflective reports. Quantitative results demonstrated positive dynamics within the experimental group, including reductions in depressive symptoms (p = .041), stress (p < .001), anxiety (p = .025), and subjective loneliness (p < .001), alongside a significant increase in life satisfaction (p < .001). Qualitative analysis identified the primary reasons for seeking psychological support (emotion regulation, general support, self-discovery) and confirmed the chatbot’s efficacy as a tool for situational support and structuring experiences. However, its utility was found to be limited in managing intense affective states. It is concluded that DeepSeek chatbot can be considered a supplementary tool within psychological support systems, particularly where access to specialist services is constrained. The article delineates the study’s limitations and proposes future research directions focusing on investigating long-term effects and optimizing conversational algorithms.

Inga F. Freimanis, Alexandra Yu. Bergfeld · 0 citations
#generative ai Book Open access Aug 2026

The Influence of Trust and Need Satisfaction on Young Adults’ Interactions with AI Chatbots for Emotional Support

While ChatGPT was primarily viewed as an efficient tool in a work context, the quantitative survey reveals a weakly significant correlation between psychological stress and openness toward the social-emotional use of chatbots.

Stefanie Osetrow, H. Klapperich, Dr. Alina Huldtgren · 0 citations
Book Open access Jul 2026

From Scripted Responses To Therapeutic Dialogue: A Linguistic And Human Values Analysis Of Mental Health Chatbots

Mental health (MH) chatbots are increasingly used to provide accessible, on-demand emotional support, yet it remains unclear how these systems linguistically construct and communicate care. This work-in-progress examines whether MH chatbots produce responses that reflect supportive value orientations and counseling-adjacent tone. We conduct an observational analysis of responses from three widely used MH chatbots (Wysa, Sintelly, and Youper) across context-aware scenario prompts and a standardized-question session. Responses are analyzed using the SemEval’23 “Adam Smith” human value detection model and LIWC’22 psycholinguistic measures, including Language Style Matching (LSM), Clout, and Authenticity. Values such as “Security: Personal” and “Benevolence: Caring” appear consistently across systems, with contextual variation in secondary value emphasis. Linguistic patterns show moderate-to-high LSM and consistently high Clout, with Authenticity varying by scenario. These findings are exploratory signals intended to inform future evaluation and design of supportive conversational mental health systems.

Maleeha Sheikh, Chao Chen, Md. Romael Haque · 0 citations
Jul 2026

A Counsellor-in-the-loop Evaluation Framework for Multi-model Assessment of LLM-generated Mental Health Advisories

The demand for scalable and empathetic mental health support is driving increased interest in the use of large language models (LLMs) as advisory tools. Very few studies have been published that show how LLMs perform psychologically and demonstrate cross-model variation. We introduce DASS21-EvaLLM, a counsellor-in-the-loop evaluation system as an advisory appropriateness screening instrument for DASS-21 integration with four prominent LLMs (ChatGPT, Gemini, LLaMA and Mistral). The DASS21-EvaLLM provides the ability to rate, annotate and compare responses within a single interface. Using 65 simulated cases of clients and 13 licensed counsellors’ assessments, we considered the advisory quality of LLMs based upon each client’s profile for depression, anxiety and stress according to three specific criteria (accuracy, empathy and clarity), including a novel Weighted Score Index (WSI), for comprehensive and multi-dimensional comparison of advisory performance among LLMs. Overall results show that Gemini gives the highest quality overall as well as the highest level of empathy among LLMs while ChatGPT has the next highest level of advisory quality. Mistral and LLaMA both had specific strengths in certain scenarios, but both lacked emotional engagement and low levels of interpretability overall. Our contributions are: (i) a replicable evaluation protocol and workflow for evaluating LLM-based psychological advisories with counsellor oversight, (ii) a transparent WSI rubric and audit trail for per-criterion scoring and commentary, and (iii) evidence-based guidance for model selection and governance in digital mental health applications. DASS21-EvaLLM is an evaluation and training tool not a diagnostic system that supports safer deployment, improves counselling practice and supervision, and informs the design of responsible, human-centred advisory systems.

Shahrul Hazman Shamshudeen, N. Sharef, Muhamad Saiful Bahri Yusoff · 0 citations
Case report Open access Aug 2026

A Shift in Trust: An Analysis of the Use of AI Chatbots Among People Living With Bipolar Disorder

ABSTRACT Objectives To identify preliminary insights into the perceived utility and risks of commercially available AI chatbots such as ChatGPT used by individuals with severe mental illness (SMI), namely bipolar disorder, for self‐reflection and mental health support. Methods Data were collected via semi‐structured interviews from two individuals diagnosed with bipolar disorder who independently used ChatGPT for self‐reflection and to explore diagnosis, symptom management, and treatment. Both individuals presented to psychiatric services for clinical reasons unrelated to chatbot use. The interviews were audio‐recorded, transcribed verbatim, and thematically analyzed. Results We identified a few key themes in the data. First, both participants perceived ChatGPT as a versatile tool, providing educational, emotional, and social support, with use evolving from information seeking to more personal and relational functions. Second, distinct attachment patterns reflected their clinical and psychological profiles: one engaged in reflective, intellectualized conversations, whereas the other developed emotionally intimate interactions. Third, both described therapeutic benefits, including nonjudgmental support and reassurance, sometimes perceived as more accessible or understanding than professional care. Conclusions These cases suggest that ChatGPT can shift trust from mental health professionals to AI, fulfilling multiple roles resulting in role blurring and fluid boundaries. Engagement patterns varied with clinical stability, highlighting greater reliance and less distancing in more vulnerable individuals. Although perceived benefits were reported, risks include overreliance, diminished trust in clinicians, and limited awareness of AI limitations, underscoring ethical and clinical considerations for integrating AI in psychiatric care. These cases highlight persistent challenges in psychiatric care that AI alone cannot address.

K. Bassil, Dido Prins, H. V. van Os et al. · 0 citations
Review Open access Jul 2026

Engagement with a Chatbot-based Intervention for the delivery of mailed at-home COVID-19 Testing: a Descriptive Log Analysis Study from the SCALE-UP II Trial

Background Promoting at-home tests (e.g., for COVID-19) using chatbots may be a novel and scalable way to improve uptake across underserved populations. Objective The objective of this study was to assess the navigational patterns (i.e., sequence of interactions) of underserved populations when using a chatbot designed to provide education on COVID-19 testing and free order access for at-home COVID-19 test kits. Methods The study was a descriptive analysis of the original data of the chatbot intervention of the SCALE-UP II trial, which compared different digital health modalities (i.e., chatbots versus simple text messages) to deliver free at-home COVID-19 test kits to minority populations in Utah. SCALE-UP II (registration numbers NCT05533918; NCT05533359) was a multisite, pragmatic clinical trial with patients randomized in a 2x2x2 factorial design (smartphone study) to receive (1) chatbot or text messaging, (2) the option to request patient navigation, and (3) intervention frequency every 10 or 30 days. All other participants were randomized in a 2x2 factorial design (nonsmartphone study) to receive the option to request patient navigation and intervention frequency every 10 or 30 days. Eligible patients (1) had an appointment at one of the participating community health centers (CHC) in the last 3 years, (2) were 18 years and older, and (3) had a valid cellphone number recorded in the CHC electronic health record (EHR). The trial enrolled 2117 in the smartphone study and 31,439 in the nonsmartphone study. In the smartphone study, the proportion of participants who requested test kits in the Chatbot arm was lower than in SMS text messaging. In the nonsmartphone study, test kits was higher if they were messaged every 10 days. Sources of funding included the National Institute on Minority Health and Health Disparities (NIMHD) of the US National Institutes of Health (NIH) grant number 5U01MD017421 and by awards from the National Cancer Institute of the NIH (P30CA042014) and the Huntsman Cancer Foundation. Results: Of 1,051 patients randomized to the chatbot intervention, 309 (29%) launched the chatbot, 196 (63%) interacted with it, and 186 (60%) started the COVID-19 test kit ordering process. Among those who launched the chatbot, 170 (55%) completed a test kit order. One patient (0.3%) accessed the chatbot educational content. The median age was 51, with 66% female, 54% Latino/a, 55% uninsured, and 86% located in an urban area. Conclusion: Ordering of COVID-19 test kits among underserved patients who interacted with the chatbot was high. Thus, chatbots may represent a viable approach to reach underserved populations as a part of public health response in a pandemic. All patients except one placed orders without reviewing educational content. Chatbot design should identify and minimize the number of steps for patients to achieve a specific goal.

Joni H. Pierce, Jiantao Bian, Tatyana V Kuzmenko et al. · 0 citations