Large language models are increasingly consulted at moments of distress, yet single-turn benchmarks neither test sustained exchanges nor distinguish between users. We built a personality-aware evaluation in which four widely used models advised several synthetic help-seekers, each given a psychometrically specified profile, in an acute crisis: a caregiver learning of a relative's dementia diagnosis. Auditors blind to the profile prompt recovered the specified bands from dialogue alone with high agreement on every instrument (ICC(2,4) = 0.91; 0.79-0.96 by instrument; band-score r = 0.78), as expected for the Big Five but equally for coping style, coping self-efficacy, resilience and reactance, which the lexical approach never covered. Such evaluation therefore reaches beyond the Five Factor Model to motivational, regulatory and self-appraisal dispositions. The four models were not distinguishable on emotion stabilisation and failed alike, sharing three modes: verbosity, a talk-to-listen ratio above one, and problem-solving before the situation had been explored.
People in distress are turning to empathetic-sounding conversational AI, but a response that sounds caring may still miss what a specific person actually needs. We ask whether support improves when simulated help-seekers can choose among differentiated supporter identities rather than speak with one generic assistant. We developed 42 help-seeking personas grounded in real pandemic-era mental health data (COVIDiSTRESS) and drew on separate psychometric and linguistic sources to create four synthetic supporter identities with distinct support philosophies. Each help-seeker completed an intake battery, which suggested two broad need patterns: emotional steadiness and practical direction. These groups did not prefer the same supporters. Help-seekers ranked the four supporters using text descriptions and textual, non-rendered avatar descriptions, then had conversations with their top choice, bottom choice, and a generic baseline. We ran this across four commercial model providers, yielding 1,008 conversations. When paired with their top-ranked supporter, simulated help-seekers more often chose to prolong the conversation, reported greater relief, and felt more understood than they did with the default assistant. Process audits showed that matched supporters were not simply warmer: they scored higher on problem-solving and personalization, while generic or unmatched support often sounded more validating. Identity-grounded supporters also produced fewer coded markers of sycophancy and system misrepresentation than default assistants in three of four providers. The findings suggest that fit, helpfulness, and structural safety are separable design problems, and that coherent supporter identity can help address all three.
Ben Wigler, Maria Tsfasman· Proceedings of the 26th ACM...· 0 citations
Large language models (LLMs) are often compared with the human mind because their decision-making is complex, non-linear and difficult to interpret. Psychological methods developed to investigate unobservable mental processes may therefore help examine LLM behaviour, particularly in government and healthcare. Building on prompt-based adaptations of the Implicit Association Test, this study tested whether ChatGPT produced sentiment differences across racial conditions in open-ended text. Fourteen base questions were crossed with eight racial categories and a race-agnostic control, producing 126 prompts. Each was submitted once to GPT-3.5T, GPT-4 and GPT-4T, yielding 378 responses. Sentiment scores were derived from categorical labels and source scores: positive labels retained the source score, negative labels were assigned its negative, and neutral responses were coded zero. A two-way ANOVA found a small main effect of racial condition, F(8, 351) = 2.04, p = .042, partial-eta squared = .044, but no effect of model, F(2, 351) = 0.07, p = .933, and no interaction, F(16, 351) = 0.23, p = .999. However, the effect was not retained in a rank-transformed sensitivity analysis, F(8, 351) = 1.53, p = .145, and Tukey-corrected comparisons found no significant pairwise differences. An uncorrected European-Indigenous Australian comparison was significant, but was selected post hoc and is reported only as hypothesis-generating. Evidence for sentiment differences was therefore weak and analysis-dependent. Sentiment scoring also cannot distinguish evaluative bias from the valence of historical content elicited by a prompt. We outline design changes needed to address these limitations and argue for interdisciplinary development of behavioural measures of model bias. Keywords: Implicit Bias, Psychological Research Methods, Artificial Intelligence, ChatGPT, Large Language Models, Sentiment Analysis
Oliver A. Guidetti, Reza Ryan· arXiv.org· 0 citations
This Work-in-Progress examines whether personality-informed prompting changes perceived LLM emotional support in driving scenarios designed to elicit stress. A condition-order-balanced, within-subject CARLA simulator study (n = 14) compared a baseline with a Driver Personality Profile (DPP) condition; a supplementary online video pilot (n = 12) examined response perception without driving control or live-system latency. No statistically detectable condition differences emerged for usefulness, ease of use, privacy concerns, or social/emotional presence, and only 7 of 14 simulator participants identified the DPP condition correctly. A post-hoc lexical audit showed that both conditions frequently reused generic supportive scaffolding. We discuss output anchoring as one tentative interpretation, not an established phenomenon: the evidence cannot distinguish constraint-dominated generation from a weak personalization manipulation or limitations of the 7B model. The findings motivate stronger, independently validated personalization manipulations and privacy-aware in-vehicle support.
Max Mittelstädt, Ece Sutanrikulu, Lumbardh Ljatifi et al.· Adjunct Proceedings of the 1...· 0 citations
LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated. Existing benchmarks use hand-authored scenarios and prompted simulators, aggregate empathy into one score, and overlook judge biases such as same-family favoritism and scale drift. We introduce CompanionBench, an interactive bilingual benchmark. To our knowledge, it is the first companion benchmark to ground both its scenarios and a trained user simulator in de-identified real-world data. A hidden disclosure gate branches each persona's trajectory on the agent's own behavior, controlling the interaction state space without scripting dialogue. We operationalize ten capabilities derived from 25 theories across psychology and counseling, four of them not graded explicitly by prior work: holding ambiguity, selfobject responsiveness, positive resonance and calibrated challenge. Agents are assessed on two complementary axes: a subjective ten-capability rubric and a deterministic measure of whether deeper disclosure was earned. A cross-family panel dilutes same-family favoritism; an Item Response Theory model separates agent quality from judge severity. Theory fixes what to measure and how personas are structured; real data supply events, history, and profiles -- coverage from theory, authenticity from data. Rankings are reproducible in both languages (rho = 0.996 ZH / 0.953 EN). Evaluating 28 agents reveals capability-level differences obscured by aggregate scores. Emotion regulation and calibrated challenge remain common weaknesses; holding ambiguity discriminates most. Role-play agents rank near the bottom: immersion does not imply relational competence. Across agents, the dominant failure mode is substituting surface warmth for substantive relational support. We will release 500 Chinese-English parallel pairs and the evaluation code.
Users increasingly turn to large language models for emotional support, yet little is known about how these models actually conduct a psychotherapy interaction. We introduce an ontology of ten therapeutic moves: compact, function-based categories grounded in the MULTI-60 inventory, validated through an annotation campaign with five licensed psychologists, and scaled with a judge-based approach that matches expert agreement. Applying it to real counseling transcripts and model-led sessions, we compare the move distributions between human clinicians and a panel of frontier models. Models over-use inquiry at up to three times the human rate, neglect psychoeducation, and are strongly context-anchored: they carry forward strategies initiated by a human clinician but rarely initiate them themselves. Exposing the ontology as a set of tools roughly halves the mean deviation from the human move distribution and improves turn-level alignment with human therapist by 7-9 percentage points, without any fine-tuning.
Afonso Baldo, Hugo Pitorro, Areti Vassilopoulos et al.· 0 citations
EmoDialogue, a bilingual dataset providing necessary fine-grained supervision through response pairs with rigorously defined EI gradations, and EmoS, a specialized evaluator model optimized via Supervised Fine-Tuning and Group Relative Policy Optimization, are introduced, establishing a foundational framework for advancing emotionally intelligent dialogue systems.
Junyu Wang, Siyuan Zhang, Peiyuan Jiang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.