Aug 2026· International Journal of Medical Informatics· Vol 220, pp.
106648
· 0 citations· 35 references
Medicine
TL;DR
LLMs hold substantial potential to enhance healthcare teamwork by supporting clinical decisions, streamlining administrative workflows, and improving patient communication, however, ethical, legal, and accountability concerns remain.
Abstract
Background
Generative artificial intelligence, particularly large language models (LLMs), has rapidly advanced and shows promise in healthcare for supporting teams through their ability to understand and generate medical text. While human-AI collaboration has been explored, the integration of LLMs into healthcare teams remains under-researched.
Objective
This scoping review aims to examine how LLMs are currently used to support teamwork and communication in healthcare teams, including both solely professional teams and those involving patients.
Methods
Following PRISMA-ScR guidelines, we registered our review with the Open Science Framework (July 30, 2025). We searched PubMed, Web of Science, and ScienceDirect for articles from 2014 to 2024. After screening 3,865 unique titles and abstracts, 127 full texts were reviewed.
Results
Twenty studies were included, predominantly employing quantitative and simulation-based designs, with limited in situ evaluations. LLM use cases were categorized into decision support, communication, and administrative functions. Outcome measures primarily focused on accuracy and quality (15/20 studies), with fewer assessing safety (4/20), readability or empathy (5/20), workflow efficiency (3/20), and error modes (2/20). Across use cases, LLMs demonstrated potential to improve efficiency and communication, although performance and risks varied by task complexity and use context.
Conclusion
LLMs hold substantial potential to enhance healthcare teamwork by supporting clinical decisions, streamlining administrative workflows, and improving patient communication. However, ethical, legal, and accountability concerns remain. Current studies largely evaluate model performance without considering the dynamics of human team members. Future research should examine LLMs' impact on trust, collaboration, and decision-making within clinical teams, while implementation efforts must address contextual and interdisciplinary factors to ensure responsible integration.
Background Large language models (LLMs) are increasingly used in health care by nonprofessionals (ie, individuals without formal training in health-related professions). These applications must be evaluated in an appropriate manner to prevent misinformation and harmful decisions. To date, guidance to evaluate LLM-based applications for nonprofessional users remains limited and fragmented, leaving researchers and developers without a scientifically grounded set of quality dimensions, metrics, and measurement tools to guide them. Objective This protocol outlines a scoping review that maps approaches for evaluation of LLM-based applications used for health purposes by nonprofessionals. It identifies current methods and maps them thematically by assigning them to evaluation dimensions, metrics, and measurement instruments. The review will provide a comprehensive overview of evaluation methods currently in use. Methods The study follows the Joana Briggs Institute approach for conducting scoping reviews and reports. The protocol is reported in accordance with the PRISMA-P (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Protocols) guidelines, and the scoping review will be reported in accordance with the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines. The inclusion criteria comprise studies that evaluate LLM-based applications that are used in the context of health care by nonprofessionals. The search was conducted in PubMed, CINAHL, PsycInfo, and IEEE Xplore. Results since 2021 were considered. Data will be summarized and interpreted qualitatively. Publication screening was conducted by 2 independent reviewers in a blinded manner, with discrepancies settled through discussion. Data extraction and charting will be performed by 1 reviewer. To ensure quality, a random 10% sample of the publications will be independently charted by a second reviewer. Disagreements in the double-extracted subset will be resolved through discussion. Results As of July 2026, a steering committee of 6 researchers has been chosen for the conduct of the review. An initial search resulted in 8538 records after removing duplicates. After screening of these 8538 publications, 17.8% (1524/8538) were eligible for retrieval, of which 88.3% (1345/1524) were retrieved. Full-text screening (completed by 1 reviewer) excluded publications due to nonmatching populations (155/1345, 11.5%), concepts (246/1345, 18.3%), and contexts (24/1345, 1.8%), as well as secondary work (14/1345, 1%), leaving 67.4% (906/1345) of these publications for data extraction. We plan to perform final full-text screening, data extraction, coding, and synthesis of results in the fourth quarter of 2026. Conclusions The scoping review aims to identify and map current evaluation methods for LLM-based applications used in health care by nonprofessionals. It will provide a systematic overview of the current state of research and insights into quality dimensions, metrics, and measurement instruments. The findings will provide directional guidance for further research and development in the field of quality assurance for LLM-based applications used by nonprofessionals. International Registered Report Identifier (IRRID) DERR1-10.2196/93509
Maren Keuchel, Pinar Bisgin, Tom Strube et al.· JMIR Research Protocols· 0 citations
Language barriers hinder healthcare, particularly during case history-taking, a key part of diagnosis. While multilingual artificial intelligence (AI) chatbots offer solutions, there is fragmented evidence of their effectiveness and impact. This systematic review followed PRISMA 2020 guidelines, examining studies published between 2015 and 2025 on multilingual AI chatbots in healthcare across four databases (Google Scholar, Scopus, Web of Science, and PubMed), using a two-stage screening process. Data extraction focused on applications, supported languages, underlying technologies, target populations, and clinical outcomes. From 503 records, 49 studies, covering primary care, telemedicine, oncology, mental health, and other areas, met the criteria. Supported languages included English, Spanish, Arabic, Chinese, Hindi, and other underrepresented languages. In individual system evaluations using heterogeneous methodologies and evaluation settings, AI chatbots achieved a diagnostic accuracy ranging from 72–92%. Core technologies included large language models (LLMs), bidirectional encoder representations from transformers (BERT), a generative pre-trained transformer (GPT), retrieval-augmented generation (RAG), speech recognition, and distillation. The findings show that these improve clinical workflow (30–70% time savings) and patient engagement, reduce language barriers, and promote health equity. However, the overall evidence certainty was low to moderate, reflecting the predominance of prototype and proof-of-concept studies. Multilingual AI chatbots demonstrate a boost in healthcare efficiency, a reduction in language barriers, and the promotion of health equity, but exhibit challenges regarding validation, workflow integration, and evaluation standards, along with ethical issues such as privacy and bias. Future research should include real-world studies, diverse populations, standardized outcome measures, and long-term equity assessments.
R. Sharanesha, Deepti Virupakshappa, A. Abushanan et al.· Informatics· 0 citations
It is suggested that the reliability of health care LLM reliability is difficult to evaluate adequately using a single universal standard, and future evaluations of health care LLM reliability need to be guided by standardized evaluation frameworks that reflect domain-specific contexts.
Euijun Yang, S. Ko, Hyekyung Woo· Journal of Medical Internet...· 0 citations
Recent developments in large language models (LLMs) have created new opportunities to support primary care, where much of clinical work is text-mediated. This narrative review synthesizes evidence on LLM applications relevant to primary care workflows and summarizes implementation safeguards. Across studies, the most consistently supported near-term value is workflow augmentation, particularly documentation and inbox management (e.g., drafting portal replies and summarizing information for clinician review) and communication support, where benefits are reported primarily as process endpoints (time, acceptability, perceived communication quality) rather than hard patient outcomes. Evidence for improvements in clinician diagnostic reasoning, treatment planning, and downstream patient outcomes is more limited and context-dependent, and many evaluations remain simulated or conducted in adjacent settings, limiting generalizability to routine primary care. Accordingly, potential roles in population health and cost reduction should be treated as hypothesis-generating and evaluated prospectively. Challenges related to privacy, security, transparency, and model reliability shape organizational governance requirements and evolving regulatory expectations for the clinical use of generative AI in primary care. We emphasize a pragmatic adoption approach: prioritize high-volume, lower-risk clerical and communication workflows; maintain clinician verification and accountability; and apply governance and equity safeguards (e.g. privacy, security, transparency, auditability, monitoring for drift and error) before scale-up. Christof et al. provide a narrative review that synthesizes evidence on LLM applications relevant to primary care workflows and summarizes implementation safeguards. They highlight the remaining need to demonstrate improved patient outcome of clinical improvement in many studies and outline a pragmatic adoption approach in practice.
Michael Christof, Krish Patel, Jiandong Zhou et al.· Communications Medicine· 0 citations
Provider-to-provider patient handoffs are a routine yet complex component of intensive care unit (ICU) workflows and are essential for patient safety. Emerging generative artificial intelligence (AI), particularly large language models (LLMs), may improve clinical communication through automated summarisation, documentation and handoff-related workflows. This scoping review followed PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines and searched six databases (PubMed, IEEE Xplore, Google Scholar, ACM Digital Library, arXiv and medRxiv) on 7 October 2025 for studies published between 14 August 1999 and 6 June 2025. Eligible publications examined LLMs or related language technologies in clinical summarisation, documentation, communication workflows, clinical decision support or patient handoffs. Forty-six studies met the inclusion criteria. Most evaluated LLMs for clinical summarisation, information extraction or decision support, whereas studies directly evaluating LLM-generated patient handoffs were rare. Only one study examined emergency department handoffs and no studies evaluated ICU-specific handoffs. Across the literature, LLMs demonstrated promising performance for clinical summarisation but recurrent challenges included hallucinations, clinically relevant omissions, limited model transparency and reliance on linguistic evaluation metrics rather than patient-centred outcomes. Human expert review was frequently incorporated to assess clinical accuracy and safety. Overall, current evidence suggests that LLMs have considerable potential to support future ICU handoff workflows but substantial evidence gaps remain. Prospective evaluation in real-world ICU settings, with clinically meaningful safety metrics and structured communication frameworks, is needed before widespread implementation.
S. Bains, Michael R Kolesnikov, Steven Bedrick et al.· Considerations in Medicine· 0 citations
AIM
To examine the overall performance of large language models (LLMs) in generating nursing care plans, clarify their role in nursing practice and identify directions for future research.
DESIGN
This study conducted a scoping review in accordance with Arksey and O'Malley's methodological framework.
METHODS
Five electronic databases were systematically searched: Web of Science Core Collection, PubMed, Scopus, CINAHL and IEEE Xplore. The search was limited to studies published between 1 June 2018 and 5 April 2026.
RESULTS
Fifteen studies were included. Existing studies primarily used nonreal patient cases and evaluated the textual quality of model-generated nursing care plans across a range of specialties. None examined LLM use within real-world clinical nursing workflows. Evaluation criteria mainly focused on accuracy, information quality and reliability, and readability. The strengths of LLMs in nursing care planning were concentrated in text organization, standardized terminology matching, and the initial drafting of nursing goals and interventions. However, important challenges remain, including privacy, hallucination, and bias.
CONCLUSIONS
LLMs may serve as assistive tools for generating initial drafts of nursing care plans, but they cannot yet replace nurses' clinical judgement. Future research should further refine evaluation frameworks and examine the impact of LLM-generated nursing care plans within real-world nursing workflows. Nurse-led human-AI collaboration should be emphasized to support the responsible translation of LLM-assisted nursing care planning into practice.
IMPACT
This scoping review highlights that, at present, LLMs can only serve as assistive tools in the development of nursing care plans, while nurses remain the primary decision-makers. It also underscores the need to enhance nurses' AI literacy to strengthen human-AI collaboration and facilitate the integration of LLMs as valuable supportive tools in nursing practice.
PATIENT OR PUBLIC CONTRIBUTION
No patient or public contribution.
Jianwen Zeng, Xule Zhu, Shiying Shen et al.· Journal of Clinical Nursing· 0 citations