Jan 2026· Journal of Nursing Management· Vol 2026· 0 citations· 90 references
Medicine
TL;DR
The findings provide nursing managers with a structured reference for identifying core evaluation dimensions, interpreting evidence, and planning the evaluation and deployment of nursing‐specific LLM applications.
Abstract
Aim To examine the current state of large language model (LLM) evaluation in clinical nursing, synthesize the evaluation dimensions and metrics reported in the literature, and develop a preliminary evaluation framework for LLMs in clinical nursing through expert consultation. Background The advancement of artificial intelligence products, exemplified by LLMs, has generated excitement about their potential applications in clinical nursing practice, but their effectiveness remains uncertain. Methods A scoping review was conducted in accordance with Arksey and O’Malley’s framework and incorporated experts’ consultation. A literature search was conducted across Web of Science, PubMed, Embase, and the Cochrane Library, from their inception to June 21, 2025. Three expert meetings involving eight experts were conducted between August and October 2025 to synthesize evaluation frameworks and scenarios. Results A total of 42 studies were included, and the GPT family was the most frequently evaluated. Thirty‐seven evaluation metrics were extracted and refined through expert consultation into six primary domains: performance and accuracy, clinical validity and safety, workflow integration and efficiency, usability and user experience, model reliability and ethical considerations, and competency development. A scenario classification and a proposed minimum technical reporting checklist were also developed to support the transparent and comparable evaluation of LLMs in clinical nursing. Conclusion This study used a scoping review and expert consultation to summarize contemporary literature on LLM evaluations in clinical nursing practice. It provides a preliminary, structured basis for developing and refining a standardized evaluation framework. Implications for Nursing Management This study highlights the need for systematic and context‐sensitive evaluation of LLMs in clinical nursing. The findings provide nursing managers with a structured reference for identifying core evaluation dimensions, interpreting evidence, and planning the evaluation and deployment of nursing‐specific LLM applications.
It is found that the region's nursing profession is undergoing a profound transformation from "technical operational labor" to "knowledge-intensive profession" and may provide references for educational system reform, clinical training, and career path planning and policy formulation.
Jun-Lin Yang, Na Wang, Bi Guan et al.· International Nursing Review· 0 citations
ChatGPT and related LLMs are best positioned as auxiliary tools that augment rather than replace professional nursing judgement, and institutions should prioritize closed-loop, enterprise-grade AI deployments over public platforms to ensure GDPR compliance.
Izabella Uchmanowicz, Heba M. Aldossary, Christopher S. Lee et al.· Journal of Clinical Nursing· 0 citations
Aim To synthesize evidence on strategies and contextual factors that influence confidence development among newly graduated nurses in acute care settings and to propose an integrated conceptual framework to guide practice, policy, and research. Design Integrative literature review. Methods This integrative review follo...
Justin Fontenot, Fu-Qin Liu, J. Immanuel· Journal of Nursing Managemen...· 0 citations
AIM
To systematically identify and categorize the outcomes, evaluation metrics, and measurement tools employed in studies evaluating the application of large language models (LLMs) in the nursing field.
BACKGROUND
LLMs are increasingly applied in nursing education, clinical practice, and research. However, outcomes,...
Zhi-Wei Lin, Xiao-Y. Huang, Yan-Hong Yan et al.· International Nursing Review· 0 citations
Background Clinical reasoning (CR) is a complex, context‐dependent process that is essential for safe nursing care. Existing reviews have primarily focused on nursing students rather than practicing clinical nurses. Rapid changes in patient acuity and the increasing integration of digital technologies into healthcare h...
It is suggested that the reliability of health care LLM reliability is difficult to evaluate adequately using a single universal standard, and future evaluations of health care LLM reliability need to be guided by standardized evaluation frameworks that reflect domain-specific contexts.
Euijun Yang, S. Ko, Hyekyung Woo· Journal of Medical Internet...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.