Large language models in healthcare: applications, evaluation frameworks, and governance pathways — a scoping review and multidimensional framework
Abstract
Large language models (LLMs) are increasingly evaluated for healthcare applications spanning clinical documentation, decision support, patient communication, research assistance, and operational workflows. Despite rapid adoption interest, evidence remains heterogeneous and standardised approaches for evaluation and governance are not yet consistently applied. To map healthcare applications of LLMs, synthesise reported outcomes and risks, and propose a multidimensional evaluation and governance framework oriented toward digital-health implementation. A scoping review was conducted following the methodological guidance of Arksey and O'Malley, Levac et al., and the Joanna Briggs Institute, with reporting aligned to the PRISMA-ScR checklist. The protocol was prospectively registered on the Open Science Framework https://doi.org/10.17605/OSF.IO/SWP78 ). Searches were performed in PubMed/MEDLINE, Embase, Scopus, Web of Science, IEEE Xplore, ACM Digital Library, and the ACL Anthology, supplemented by medRxiv/bioRxiv preprint searches and grey-literature scanning of WHO, FDA, and EU AI-Office guidance, covering 1 January 2019–30 April 2026. Title/abstract and full-text screening were carried out in duplicate; inter-rater agreement was Cohen's κ = 0.78. Data were extracted in duplicate using a piloted form. Synthesis followed Braun and Clarke's six-phase reflexive thematic analysis. The complete list of 78 included studies is provided in Supplementary File S4. Seventy-eight studies met inclusion criteria; 68% were published between 2023 and 2025. Five application domains were identified: (1) clinical documentation and summarisation; (2) clinical decision support and reasoning assistance; (3) patient communication and health-literacy support; (4) biomedical research and knowledge synthesis; and (5) administrative and operational use cases. The evidence base was dominated by benchmark and simulated-workflow studies (89%), with limited prospective workflow-embedded evaluations (11%). Reported benefits concentrated on documentation efficiency, text quality, and knowledge synthesis; safety-relevant risks included hallucinated content, omission of clinically critical information, demographic bias, privacy vulnerabilities, limited explainability, and automation bias. Studies were geographically concentrated in North America and East Asia, with limited representation from Sub-Saharan Africa, South Asia, and Latin America. Current evidence supports cautious deployment of LLMs in selected healthcare tasks under structured oversight. Translational progress depends on prospective evaluation, standardised reporting, equity-focused audits, and lifecycle governance with continuous monitoring. The proposed five-dimensional framework (technical performance, clinical validity, equity, workflow integration, governance) coupled with a three-tier risk model is intended to support researchers and healthcare organisations in assessing readiness and implementing LLM-enabled tools responsibly.