Users increasingly turn to large language models for emotional support, yet little is known about how these models actually conduct a psychotherapy interaction. We introduce an ontology of ten therapeutic moves: compact, function-based categories grounded in the MULTI-60 inventory, validated through an annotation campaign with five licensed psychologists, and scaled with a judge-based approach that matches expert agreement. Applying it to real counseling transcripts and model-led sessions, we compare the move distributions between human clinicians and a panel of frontier models. Models over-use inquiry at up to three times the human rate, neglect psychoeducation, and are strongly context-anchored: they carry forward strategies initiated by a human clinician but rarely initiate them themselves. Exposing the ontology as a set of tools roughly halves the mean deviation from the human move distribution and improves turn-level alignment with human therapist by 7-9 percentage points, without any fine-tuning.
Large language models are increasingly consulted at moments of distress, yet single-turn benchmarks neither test sustained exchanges nor distinguish between users. We built a personality-aware evaluation in which four widely used models advised several synthetic help-seekers, each given a psychometrically specified profile, in an acute crisis: a caregiver learning of a relative's dementia diagnosis. Auditors blind to the profile prompt recovered the specified bands from dialogue alone with high agreement on every instrument (ICC(2,4) = 0.91; 0.79-0.96 by instrument; band-score r = 0.78), as expected for the Big Five but equally for coping style, coping self-efficacy, resilience and reactance, which the lexical approach never covered. Such evaluation therefore reaches beyond the Five Factor Model to motivational, regulatory and self-appraisal dispositions. The four models were not distinguishable on emotion stabilisation and failed alike, sharing three modes: verbosity, a talk-to-listen ratio above one, and problem-solving before the situation had been explored.
P. Fonseca, R. Rodríguez-Carvajal, Rafael A. Calvo· 0 citations
LLM-based simulated clients are increasingly used to train novice counselors, evaluate LLM therapists, and generate synthetic data. However, current simulators produce overly cooperative clients that disclose too readily, accept therapeutic reframes without resistance, and resolve core issues within a single session. We trace these issues to profiles that lack causal depth and behavioral mechanisms that treat all content as equally accessible. We present PatientAct, a framework for client simulation grounded in established clinical theories. Our profiles integrate the 5Ps clinical case formulation, providing causal depth without tying the design to any single therapeutic modality. During simulation, profiles include a dynamic memory layer in which items carry trust thresholds (e.g., symptoms are available early, whereas formative memories require a sustained therapeutic alliance). At each turn, the client's emotional reaction and behavior are modeled before generating a response. If the therapist approaches gated content, PatientAct expresses resistance in terms of quantity, content, and style rather than defaulting to cooperation or a single resistance pattern. We evaluate our framework on 40 clinical situations and demonstrate that it generates diverse profiles with high clinical plausibility. Moreover, PatientAct significantly outperforms the baselines, yielding substantial gains in resistance quality and behavioral realism. Our code and data are publicly available via github.com/Sahandfer/PatientHub.
Sahand Sabour, NG TszYam, Yaqian Chen et al.· 1 citation· ⚡1
How does change begin in therapy? This article suggests that each theoretical perspective reveals only part of the whole. Change often arises not only through insight, but also when something small shifts in everyday life. Drawing on systems thinking and the philosophy of Bruno Latour, the article seeks to show that problems are not located solely within an individual, but emerge in a network of relationships, habits, tensions, words, objects, and circumstances. From this perspective, the article introduces systemic field experiments: small, meaningful actions in daily life through which people can explore what brings about positive change. Such field experiments can help shift a stuck pattern step by step. Through a case study of a boy with obsessive-compulsive symptoms, the article shows how slight changes within a family can have far-reaching effects. Therapy thus becomes a shared process of trying, observing, and discovering what works.
The article analyses the capabilities and limitations of large language models’ (LLMs) powering chat interfaces, such as ChatGPT, in the context of their potential use as “psychotherapists”. The authors relate this issue to the principles of psychodynamic psychotherapy, in which interpersonal relationships, unconscious processes, transference, and countertransference are of key importance. Based on a review of the literature and technical analysis, it is indicated that contemporary LLMs do not meet the basic conditions for conducting therapy in this field. They lack awareness, emotionality, and the ability to mentalize, whereas their “empathy” is purely linguistic and simulated. Human interactions with the model can create the illusion of a therapeutic relationship, promote anthropomorphizing and excessive trust in the algorithm. The article highlights the ethical, legal, and purely psychological risks associated with assigning chatbots the role of therapists. The analysis encompasses problems of hallucinations, data bias, and lack of ethical oversight. At the same time, the authors point to the potential benefits of using LLMs as support tools – in psychological education, preliminary diagnosis, stress reduction, or as a support for therapists in the process of analysis. They emphasize the need for an interdisciplinary approach combining psychology, law, and ethics, as well as the need to develop standards for safety and oversight of the use of LLMs in the field of mental health. The article concludes with a call for caution in the face of the growing integration of technology into therapeutic practice.
Mateusz Łabuz, Paweł Szczęsny, Katarzyna Mika-Łabuz· Ethics and Information Techn...· 0 citations
Large language models (LLMs) are often compared with the human mind because their decision-making is complex, non-linear and difficult to interpret. Psychological methods developed to investigate unobservable mental processes may therefore help examine LLM behaviour, particularly in government and healthcare. Building on prompt-based adaptations of the Implicit Association Test, this study tested whether ChatGPT produced sentiment differences across racial conditions in open-ended text. Fourteen base questions were crossed with eight racial categories and a race-agnostic control, producing 126 prompts. Each was submitted once to GPT-3.5T, GPT-4 and GPT-4T, yielding 378 responses. Sentiment scores were derived from categorical labels and source scores: positive labels retained the source score, negative labels were assigned its negative, and neutral responses were coded zero. A two-way ANOVA found a small main effect of racial condition, F(8, 351) = 2.04, p = .042, partial-eta squared = .044, but no effect of model, F(2, 351) = 0.07, p = .933, and no interaction, F(16, 351) = 0.23, p = .999. However, the effect was not retained in a rank-transformed sensitivity analysis, F(8, 351) = 1.53, p = .145, and Tukey-corrected comparisons found no significant pairwise differences. An uncorrected European-Indigenous Australian comparison was significant, but was selected post hoc and is reported only as hypothesis-generating. Evidence for sentiment differences was therefore weak and analysis-dependent. Sentiment scoring also cannot distinguish evaluative bias from the valence of historical content elicited by a prompt. We outline design changes needed to address these limitations and argue for interdisciplinary development of behavioural measures of model bias. Keywords: Implicit Bias, Psychological Research Methods, Artificial Intelligence, ChatGPT, Large Language Models, Sentiment Analysis
Oliver A. Guidetti, Reza Ryan· arXiv.org· 0 citations
Recent advancements in large language models have revolutionized the field of psychological counseling, especially in the context of Cognitive Behavioral Therapy (CBT). While the success of CBT relies heavily on dynamic decision-making informed by the client's real-time mental state, this aspect has often been overlooked in current research, limiting both flexibility and therapeutic outcomes. In this paper, we introduce StratCBT, a dataset specifically designed for psychological counseling conversations with CBT Strategies, consisting of 9,688 sessions and around 256K utterances, with each counselor's response aligned with one of eight distinct strategies. The creation of StratCBT involves modeling clients based on their negative thoughts and generating high-quality counseling conversations through self-chat, incorporating realistic sessions as guidance, thereby significantly surpassing existing datasets in both general counseling and CBT-specific skills. We conduct extensive experiments to demonstrate the effectiveness of strategy-aligned generation and evaluate its efficacy in delivering professional and effective counseling with LLM-simulated clients to reflect real-world scenarios. The dataset can be obtained from https://github.com/zimuwangnlp/StratCBT.
Zi-Mu Wang, Yi-Wen Jiang, Xiang-Yu Zhao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.