A LINE-deployed conversational AI system grounded in the WHO iSupport for Dementia framework, built around an Assessment-Before-Intervention dialogue mechanism, which offers practical implications for the design of AI-assisted care systems in dementia and other emotionally sensitive healthcare contexts.
Abstract
Family caregivers of people with dementia face daily behavioral challenges—aggression, wandering, agitation—that require timely, contextual guidance. Existing chatbot systems typically respond with direct advice, bypassing assessment of behavioral context and caregiver emotional state. This paper presents a LINE-deployed conversational AI system grounded in the WHO iSupport for Dementia framework, built around an Assessment-Before-Intervention dialogue mechanism. The system applies a four-step reasoning process derived from the iSupport ABC behavioral cycle, governed by five safety principles, to determine when to clarify before advising. A dual-layer response model ensures emotional acknowledgment is never omitted; a hybrid keyword-semantic Retrieval-Augmented Generation (RAG) architecture bridges the lay-to-clinical vocabulary gap. We evaluated the system through a formative review with four domain reviewers in dementia care, covering 13 BPSD scenarios (52 evaluations across five quality dimensions). Mean scores (4.78–4.81 on a 5-point Likert scale) are interpreted as preliminary perceived-appropriateness data rather than clinical effectiveness evidence. The principal contribution lies in the qualitative findings: five recurrent failure modes and a structural pattern of cultural misalignment between the international iSupport framework and Taiwanese caregiving realities. Findings offer practical implications for the design of AI-assisted care systems in dementia and other emotionally sensitive healthcare contexts.
Cognitive Behavioral Therapy (CBT) provides a structured framework for understanding a user's mental state by examining the interaction between cognitive and behavioral factors. However, out-of-the-box LLMs respond fluently and empathetically, yet collapse into validation&reflection, regardless of what the user actually needs. They know theoretical CBT (scoring up to 96% accuracy on licensing exam questions) but fail to apply it effectively. We explore this gap with a knowledge-guided framework that treats CBT dialogue as controlled affective reasoning: user narratives are decomposed into Beck's Cognitive Conceptualization structure, grounded in clinical SNOMED CT concepts validated via Natural Language Inference, and a Multiple Chain-of-Thought (MCoT) strategy selection between Validation&Reflection, Socratic Questioning, or Alternative Perspectives. To measure whether such guidance actually changes behavior, we introduce the Protocol Leverage Force (F), a behavior-level metric that captures how far an intervention shifts a model away from its default response. Across three open-weight LLMs and 14 RealCBT-derived case studies, evaluated with human experts, valence-arousal trajectories, and linguistic entrainment, F shows that simply introducing protocol definitions via single chain-of-thought prompting fails to change LLM behavior, while MCoT on these definitions guides strategy selection better. Still, the effect stays within 1% (approx. 1.2-1.3%), and all models remain biased toward Validation&Reflection. These results show CBT knowledge alone does not ensure effective application, giving the affective-computing community instrumentation to measure where LLMs fall short.
Vaishnavi Sinha, Pooja Guttal, Pranay Deep Reddy Katike et al.· 0 citations
A multimodal emotion-aware architecture, which pays attention to memory-enhanced personalization and emotion-specific reinforcement learning, is introduced and hybrid human-AI approaches, which focus on safety and empathetic conversation to improve current mental health systems are recommended.
Large language models (LLMs) are expanding the capabilities of socially assistive robots (SARs) through natural dialogue, personalisation, multimodal reasoning, retained interaction context, and adaptive behaviour in healthcare. Integrating generative language models into robots, however, complicates evaluation because fluent output may exaggerate perceived competence and increase the risks of hallucination, overtrust, privacy exposure, relationship dependency, and unsafe reliance on advice or actions. This PRISMA-informed review synthesises healthcare robotics, human–robot interaction, LLM-enabled systems, ethics, implementation, and care delivery. Database searches returned 128 records, of which 110 were unique after deduplication. Supplementary retrieval and assessment yielded 85 substantive sources spanning background mapping, primary analysis, and governance. Studies focused mainly on feasibility, usability, acceptability, dialogue quality, and short-term engagement, whereas longitudinal safety, governance of retained interaction context, comparative effectiveness, workflow integration, and sustained healthcare value received limited attention. These gaps indicate that evaluation of LLM-enabled SARs must account for physical presence, social role, interaction memory, and potential actions rather than focus on conversational performance alone. The review therefore proposes HEART, a healthcare-specific evaluative architecture comprising Human-Centred Communication, Ethical and Trustworthy Deployment, Adaptive and Embodied Intelligence, Relationship Continuity, and Translational Healthcare Value. HEART uses boundary rules, operational indicators, qualitative labels, and non-additive deployment gates to separate evaluative domains, define assessable outcomes, summarise reported support, and prevent strengths in one area from masking critical safety or governance failures. Future research should validate HEART through longitudinal and comparative assessment of hallucination severity, language-to-action safety, long-term effects, equity, and post-deployment monitoring.
Conversational artificial intelligence (AI) has shown potential to support knowledge translation, personalized education, and access to health-related information. However, applications in neurodevelopmental care remain largely focused on screening and assessment, while conversational systems tailored to occupational therapy are scarce. Moreover, general-purpose generative AI may produce inaccurate, insufficiently contextualized, or clinically inappropriate responses, highlighting the need for evidence-based, domain-specific systems with robust safety mechanisms. This study describes the protocol for the co-design, development, and preliminary evaluation of a conversational AI system designed to support occupational therapists and parents of children aged 5–12 years with neurodevelopmental disorders in Greece. The system is intended as an educational and decision-support resource and will not replace diagnosis, clinical judgment, or individualized intervention. A co-designed, multi-phase, mixed-methods proof-of-concept design will be adopted, informed by the Medical Research Council framework, Design Science Research, user-centered design principles, and the CeHRes Roadmap 2.0. Development will include evidence synthesis, stakeholder needs assessment, knowledge base construction, iterative prototype development, and expert, technical, safety, and user evaluation. The system will integrate a curated occupational therapy knowledge base, retrieval-augmented generation, role-specific prompting, source verification, and layered safety guardrails. Expected outputs include a stakeholder-informed Greek-language minimum viable product and a transparent framework linking evidence, user requirements, technical design, and evaluation criteria. Preliminary evaluation will assess factual accuracy, evidence concordance, occupational therapy relevance, clinical appropriateness, safety, usability, acceptability, and perceived usefulness. This protocol provides a reproducible foundation for developing clinically relevant conversational AI in occupational therapy and for future feasibility and effectiveness studies.
Pantelis Pergantis, N. Bardis, Charalabos Skianis et al.· Brazilian Journal of Science· 0 citations
Developmental dyslexia involves persistent difficulties in word-level reading and decoding, requiring sustained linguistic practice that is difficult to maintain without supervision. Although Generative AI offers personalized support, standard Large Language Models (LLMs) often lack the pedagogical and therapeutic knowledge required for linguistic intervention. We present Foxy, a proactive LLM-driven Conversational Agent designed to assist Italian children aged 8–11 with dyslexia during morphological training. Through a modular prompt-orchestration framework and an event-driven architecture, Foxy acts as a specialized tutor that provides scaffolding, limits topic drift, and reduces hallucinations via a verified lexical knowledge base. The system was refined through expert-led co-design and evaluated in a pilot study with 18 educators, and therapists. Results indicate positive perceptions of usability, usefulness, and appropriateness, with recognition of Foxy’s motivational benefits. These findings suggest Foxy is a promising tool for dyslexia intervention, pending wider empirical validation.
Giulia Valcamonica, Giovanni Caleffi, Francesco Piferi et al.· International Conference on...· 0 citations
Social robots offer a promising means of supporting cognitive therapies for dementia care by guiding structured conversation and therapeutic activities. However, little is known about the conversational dynamics that emerge during robot-delivered cognitive stimulation therapy (CST) sessions. This study analysed the interaction patterns from robot-delivered individual CST (iCST) sessions conducted with people living with dementia in home settings. Our Co-STAR (Cognitive Stimulation Therapy by an Autonomous Robot) system was deployed in the homes of eight PwDs for one week, who completed 30-minute sessions. Conversational metrics, including words per turn, speech production rate, response duration, response latency, and self-referential language, were analysed to examine how conversational engagement is shaped by prompt personalisation, interaction phase, and participant characteristics. The findings highlight three key interactional properties of robot-delivered iCST. First, personalised prompts significantly increase response duration, self-referential language, and overall engagement compared to generic prompts. Second, conversational behaviour changes within sessions, with a reduction in the verbal output and autobiographical engagement observed during later interaction phases, which suggests cognitive fatigue. Third, first-session conversational metrics were associated with long-term participation, while living situation influenced conversational engagement patterns. These findings provide empirical insights into the factors that shape conversational engagement in robot-delivered iCST. They inform the design of adaptive conversational robots for dementia therapy.