Skip to content

Understanding Clinical Cognitive Dialogues Using Large Language Models

Sep 2026 · 0 citations · 38 references
Computer Science

TL;DR

An de-identified corpus of 33 cognitive assessment conversations with 8,250 utterances annotated for three speaker roles and 56 dialogue acts is presented, showing that broad conversational intent is easier to recognize than fine-grained communicative function.

Abstract

In-person cognitive assessment is both a test and an interaction. Clinicians explain tasks, repair misunderstandings, and adapt to patient responses, while patients may hesitate, seek clarification, or disengage. Yet clinical dialogue resources rarely label the interaction structure needed to study these behaviors at scale. We present an de-identified corpus of 33 cognitive assessment conversations with 8,250 utterances annotated for three speaker roles and 56 dialogue acts. We use this corpus to benchmark large language models on fine-grained dialogue-act classification and next-patient-utterance generation. We also test whether out-of-domain instruction data and explanation-augmented training transfer to this clinical setting. Instruction tuning produces the strongest patient-utterance reference matching and improves classification accuracy. Reasoning-aware fine-tuning produces the strongest classification results among the LLaMA-3.1-8B variants. However, even the best models struggle to separate closely related dialogue acts, showing that broad conversational intent is easier to recognize than fine-grained communicative function. The corpus and benchmark make interaction structure measurable in cognitive assessments and support follow-up work on conversational markers, clinician education, and carefully validated simulated patients. This work does not make diagnostic claims. Instead, it provides the data and evaluation framework needed to study these applications.

View source

Similar papers

TalkFa: A Unified Benchmark for Farsi Dialogue Generation and Understanding

Farsi, spoken by more than 120 million people, lacks a comprehensive benchmark for dialogue generation and understanding. We introduce TALKFA, a unified benchmark comprising three complementary datasets: (1) WIKI-FADIAL, 4.2K Wikipedia-grounded dialogues for knowledge-grounded generation; (2) DAILYDIALOG-FA, 6.6K dialo...

Neda Jamshidi, Kamyar Zeinalipour, F. Akbari et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Exposing Weaknesses in Emotion Recognition in Conversations

An LLM-as-Judge framework is introduced that evaluates each emotion independently according to its plausibility in the conversational context rather than enforcing a single-label decision, suggesting that standard single-label evaluation is therefore insufficient.

Amir Ben Khalifa, Fanny Bezancon, B. Abdulrazak et al. · 0 citations
#natural language process... Preprint Sep 2026

BaatCheet: A Multilingual Corpus for Dialogue Translation in Indian Languages

Existing translation models are typically trained on sentence-level and formal text, limiting their ability to capture everyday conversational dialogue phenomena such as informality, speaker interaction, and discourse coherence. Most existing Indic translation resources and evaluation benchmarks focus on sentence-level...

Priyanka Dasari, Yuvrajsinh Bodana, Vandan Mujadia et al. · 0 citations
Open access 2026

Retrieving Responses Useful as References in Counseling via LLM-Generated Structurally Analogous Dialogues

Retrieving past similar dialogues to assist response generation in the present context is a promising approach to addressing the shortage of highly experienced professionals in various domains, such as user support and mental health counseling. However, unlike traditional retrieval, this task requires finding dialogue...

Yu Nakagawa, Nozomu Ikeda, Kotaro Funakoshi et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.