Skip to content
Review Open access

An Evaluation Framework for Large Language Models in Clinical Nursing: A Scoping Review and Expert Consultation

Jan 2026 · Journal of Nursing Management · Vol 2026 · 0 citations · 90 references
Medicine

TL;DR

The findings provide nursing managers with a structured reference for identifying core evaluation dimensions, interpreting evidence, and planning the evaluation and deployment of nursing‐specific LLM applications.

Abstract

Aim To examine the current state of large language model (LLM) evaluation in clinical nursing, synthesize the evaluation dimensions and metrics reported in the literature, and develop a preliminary evaluation framework for LLMs in clinical nursing through expert consultation. Background The advancement of artificial intelligence products, exemplified by LLMs, has generated excitement about their potential applications in clinical nursing practice, but their effectiveness remains uncertain. Methods A scoping review was conducted in accordance with Arksey and O’Malley’s framework and incorporated experts’ consultation. A literature search was conducted across Web of Science, PubMed, Embase, and the Cochrane Library, from their inception to June 21, 2025. Three expert meetings involving eight experts were conducted between August and October 2025 to synthesize evaluation frameworks and scenarios. Results A total of 42 studies were included, and the GPT family was the most frequently evaluated. Thirty‐seven evaluation metrics were extracted and refined through expert consultation into six primary domains: performance and accuracy, clinical validity and safety, workflow integration and efficiency, usability and user experience, model reliability and ethical considerations, and competency development. A scenario classification and a proposed minimum technical reporting checklist were also developed to support the transparent and comparable evaluation of LLMs in clinical nursing. Conclusion This study used a scoping review and expert consultation to summarize contemporary literature on LLM evaluations in clinical nursing practice. It provides a preliminary, structured basis for developing and refining a standardized evaluation framework. Implications for Nursing Management This study highlights the need for systematic and context‐sensitive evaluation of LLMs in clinical nursing. The findings provide nursing managers with a structured reference for identifying core evaluation dimensions, interpreting evidence, and planning the evaluation and deployment of nursing‐specific LLM applications.

Read PDF

Similar papers

Review Open access Sep 2026

Healthcare Professionals' Perspectives on Nursing Competency: A Synthesis of Qualitative Evidence.

It is found that the region's nursing profession is undergoing a profound transformation from "technical operational labor" to "knowledge-intensive profession" and may provide references for educational system reform, clinical training, and career path planning and policy formulation.

Jun-Lin Yang, Na Wang, Bi Guan et al. · 0 citations
Review Open access Aug 2026

ChatGPT and Large Language Models in Contemporary Nursing.

ChatGPT and related LLMs are best positioned as auxiliary tools that augment rather than replace professional nursing judgement, and institutions should prioritize closed-loop, enterprise-grade AI deployments over public platforms to ensure GDPR compliance.

Izabella Uchmanowicz, Heba M. Aldossary, Christopher S. Lee et al. · 0 citations
Review Open access Jan 2026

A Multifactorial Model for Building Professional Confidence in New Graduate Nurses: Integrative Review of Influential Strategies and Contexts

Aim To synthesize evidence on strategies and contextual factors that influence confidence development among newly graduated nurses in acute care settings and to propose an integrated conceptual framework to guide practice, policy, and research. Design Integrative literature review. Methods This integrative review follo...

Justin Fontenot, Fu-Qin Liu, J. Immanuel · 0 citations
Review Open access Sep 2026

Outcomes, Evaluation Metrics, and Measurement Tools for Large Language Model Applications Research in Nursing: A Scoping Review.

AIM To systematically identify and categorize the outcomes, evaluation metrics, and measurement tools employed in studies evaluating the application of large language models (LLMs) in the nursing field. BACKGROUND LLMs are increasingly applied in nursing education, clinical practice, and research. However, outcomes,...

Zhi-Wei Lin, Xiao-Y. Huang, Yan-Hong Yan et al. · 0 citations
Review Open access Jan 2026

Factors Influencing Clinical Reasoning Among Clinical Nurses: An Integrative Review Using the Decision‐Making Ecology Model

Background Clinical reasoning (CR) is a complex, context‐dependent process that is essential for safe nursing care. Existing reviews have primarily focused on nursing students rather than practicing clinical nurses. Rapid changes in patient acuity and the increasing integration of digital technologies into healthcare h...

Ju-Hee Lee, Deokhyun Lee, Hyeran Park · 0 citations
Review Open access Aug 2026

The Reliability of Human Evaluation of Large Language Models in Health Care Settings: Scoping Review

It is suggested that the reliability of health care LLM reliability is difficult to evaluate adequately using a single universal standard, and future evaluations of health care LLM reliability need to be guided by standardized evaluation frameworks that reflect domain-specific contexts.

Euijun Yang, S. Ko, Hyekyung Woo · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.