Skip to content
Review Open access

Applications, Challenges, and Future Directions of Large Language Models in Health Care Communication: Scoping Review

Jun 2026 · Journal of Medical Internet Research · Vol 28, pp. e84726-e84726 · 0 citations · 174 references
Medicine

TL;DR

This scoping review uses communication accommodation theory to systematically map the application patterns and developmental landscape of LLM-mediated health care communication.

Abstract

Abstract Background Effective health care communication is crucial in the medical field. However, effective communication in clinical practice still faces numerous obstacles, and large language models (LLMs) offer various possibilities for improving the quality of medical communication. To date, there are no published reviews on the use of LLMs in health care communication. Objective This review sought to summarize the applications and challenges of LLMs in health care communication and to identify directions for future research. Methods A comprehensive literature search was conducted in PubMed, Embase, Web of Science, and the Cochrane Library from January 2018 to November 2025. The search and selection process followed the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guideline and the PRISMA-S (Preferred Reporting Items for Systematic Reviews and Meta-Analyses literature search extension) checklist. Eligible studies used LLMs to facilitate health care communication among the public, patients, and clinicians. Following rigorous data extraction and cross-checking, we conducted a quantitative analysis of characteristics of the included literature. Furthermore, using communication accommodation theory as a framework, we identified application patterns of LLMs in health care communication and summarized current challenges and future directions. Results Ninety-six studies were included in this review, all published between 2023 and 2025, summarizing 4 patterns of LLM application in health care communication: transforming medical information (n=30), facilitating dynamic interaction (n=38), empowering communication capabilities (n=10), and optimizing clinical workflows (n=18). The role of LLMs in health care communication is undergoing a paradigm shift from “static information processing” to “dynamic intelligent interaction.” Although they show great promise for practical applications, current evaluation methods and dimensions exhibit significant heterogeneity. Furthermore, LLMs still face multiple challenges in their practical application in health care communication, including technical reliability issues, social trust and adoption, interaction and access barriers, and clinical integration challenges. Conclusions Unlike previous studies that merely touched upon the challenges and future directions, this scoping review uses communication accommodation theory to systematically map the application patterns and developmental landscape of LLM-mediated health care communication. Health care communication powered by LLMs holds significant innovation potential and is currently still in the early stages of rapid development. Future research should focus on optimizing model performance, strengthening ethical governance frameworks, enhancing human-machine collaboration models, and ensuring responsible application of LLMs in health care through rigorous empirical validation.

Read PDF

Similar papers

Review Open access Aug 2026

The Reliability of Human Evaluation of Large Language Models in Health Care Settings: Scoping Review.

It is suggested that the reliability of health care LLM reliability is difficult to evaluate adequately using a single universal standard, and future evaluations of health care LLM reliability need to be guided by standardized evaluation frameworks that reflect domain-specific contexts.

Euijun Yang, S. Ko, Hyekyung Woo · 0 citations
Review Aug 2026

How large language models can be used for teamwork and communication in healthcare settings: A scoping review.

LLMs hold substantial potential to enhance healthcare teamwork by supporting clinical decisions, streamlining administrative workflows, and improving patient communication, however, ethical, legal, and accountability concerns remain.

Ilse Super, Olya Rezaeian, Onur Asan · 0 citations
Review Open access Aug 2026

Large Language Models for Mental Health Prediction: Scoping Review of Bias and Clinical Utility Documentation in 2019-2024

Abstract Background A growing body of literature leverages large language models (LLMs) to make mental health predictions. However, these models are prone to bias, and studies to validate their clinical utility are lacking. Objective This scoping review aims to uncover bias and clinical utility limitations stemming from the methodological design of LLM-based mental health predictive systems. In addition, it intends to document the level of self-reflection about bias and clinical challenges reported by authors in their own work. Methods This work follows the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines and was registered online. Eligible studies were original research articles in English published between 2019 and 2024, using LLMs to detect mental health conditions in nonsynthetic textual data. The search was conducted in 5 scientific databases (PubMed, Web of Science, IEEE Xplore, ACM Digital Library, and ACL Anthology) with queries associating keywords related to “Mental Health,” “Large Language Models,” and “Prediction.” We extracted both methodological information about the included studies and authors’ statements relevant to issues of bias and clinical utility. This extraction was based on a framework screening the entire pipeline of development of LLMs with applications in mental health: research design and selection, data collection, outcome definition, model development, and postdeployment considerations. Statistical description of the retrieved entities, as well as thematic coding, was performed for analysis. Results A total of 2472 articles were identified, of which 263 (10.6%) were assessed for eligibility, and 201 (8.1%) were included in the review. Included studies were mostly recent, indicating a growing interest in the use of LLMs for mental health predictions. Our analysis revealed that a majority of studies share similar methodological choices along their development pipeline: most of them focus on depressive disorders identified via processing user texts on social media, mainly with the use of nonspecialist LLMs derived from BERT (Bidirectional Encoder Representations from Transformers). Following previous works on these matters, we highlighted how these choices may hinder the clinical relevance and fairness of the envisioned systems. Similarly, we found that 164 (81.6%) studies mention themes related to bias and clinical utility; however, most of the discussion revolves around data-centered issues. Only 41 (20.4%) articles mention themes associated with at least 3 out of 5 pipeline steps, suggesting a limited appropriation of the notions of bias and clinical utility in such a sensitive context as mental health analysis. Conclusions Bias and clinical utility are lightly covered in the field of LLM-based mental health prediction research as of 2019‐2024. In-depth approaches involving interdisciplinary teams of clinicians and natural language processing specialists are needed to ensure technical soundness, clinical relevance, and fair outcomes for potential users.

Clémentine Bleuze, Karen Fort, Vincent P. Martin et al. · 0 citations
Review

Large Language Models in Medicine: Opportunities, Limitations, and Future Directions

Current evidence indicates that LLMs have substantial potential to enhance healthcare delivery, research, and personalized medicine, but they should currently be regarded as supportive tools rather than autonomous clinical decision-makers.

Antoni Klamka, Paulina Kawalec, Kamil Bronikowski et al. · 0 citations
Review Open access Aug 2026

Safety-Oriented Evaluation of Large Language Models in Health Care: Guideline-Informed Systematic Review

Abstract Background Large language models (LLMs) are rapidly emerging in health care, offering opportunities in decision support, education, and research, but raising critical concerns about safety, reliability, and ethics. Although several guidelines for trustworthy AI exist in business and technology, few systematic reviews have applied them to medical contexts. Objective This study aimed to conduct a systematic review of LLM research in health care, applying the AI Guidelines for Business as a framework across 11 domains, including safety, reliability, ethics, transparency, fairness, inclusiveness, privacy, security, robustness, data quality, and verifiability. Methods Following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 guidelines (retrospectively registered on the Open Science Framework; DOI 10.17605/OSF.IO/P4KSB), the PubMed, Scopus, Web of Science, arXiv, and IEEE Xplore databases were searched on January 15, 2025. Records were screened in 2 stages by 3 reviewers (with records retained only upon unanimous agreement). A total of 247 studies were included, of which 211 (85.4%) contributed quantitative values. Eligible studies were classified across 11 trustworthy AI domains. Heterogeneous metrics were summarized within metric families; when multiple models were evaluated, the mean across models was used as the primary estimate, with best, median, and primary-model sensitivity analyses. The LLM-assisted categorization (GPT-5 mini) was validated by using an automated internal consistency check, and 95% CIs were estimated by using cluster bootstrap on study-level values. Results Of the 25,156 records, 247 (1.0%) studies were included, and of these, 211 (85.4%) contributed quantitative values. Evaluation concentrated on accuracy (143/247, 57.9%) and fairness and inclusiveness (93/247, 37.7%), followed by data quality (47/247, 19.0%) and prevention of misinformation (44/247, 17.8%). Normalized performance was moderate to high (accuracy mean 0.73, 95% CI 0.7-0.76; data quality: 0.64; prevention of misinformation: 0.82). Selecting the best-performing model inflated domain means by up to 0.05. Privacy protection (2/247, 0.8%) and security assurance (0/247, 0.0%) were almost entirely absent. Domain assignments were recoverable from objective metric types in 98.9% of values (Cohen κ=0.985). Conclusions Current evaluations emphasize accuracy while underreporting privacy, security, robustness, explainability, and verifiability. This finding reflects gaps in reporting rather than demonstrated poor performance, underscoring the need for comprehensive, guideline-based, multidomain evaluation before deployment in high-stakes clinical settings.

Hikaru Matsuoka, Takayuki Takahashi, Takayuki Semitsu et al. · 0 citations