Skip to content
#explainable ai Review Open access

Large language models in healthcare: applications, evaluation frameworks, and governance pathways — a scoping review and multidimensional framework

Sep 2026 · Frontiers in Digital Health · 0 citations · 78 references
Artificial Intelligence in Healthcare and Education

Abstract

Large language models (LLMs) are increasingly evaluated for healthcare applications spanning clinical documentation, decision support, patient communication, research assistance, and operational workflows. Despite rapid adoption interest, evidence remains heterogeneous and standardised approaches for evaluation and governance are not yet consistently applied. To map healthcare applications of LLMs, synthesise reported outcomes and risks, and propose a multidimensional evaluation and governance framework oriented toward digital-health implementation. A scoping review was conducted following the methodological guidance of Arksey and O'Malley, Levac et al., and the Joanna Briggs Institute, with reporting aligned to the PRISMA-ScR checklist. The protocol was prospectively registered on the Open Science Framework https://doi.org/10.17605/OSF.IO/SWP78 ). Searches were performed in PubMed/MEDLINE, Embase, Scopus, Web of Science, IEEE Xplore, ACM Digital Library, and the ACL Anthology, supplemented by medRxiv/bioRxiv preprint searches and grey-literature scanning of WHO, FDA, and EU AI-Office guidance, covering 1 January 2019–30 April 2026. Title/abstract and full-text screening were carried out in duplicate; inter-rater agreement was Cohen's κ = 0.78. Data were extracted in duplicate using a piloted form. Synthesis followed Braun and Clarke's six-phase reflexive thematic analysis. The complete list of 78 included studies is provided in Supplementary File S4. Seventy-eight studies met inclusion criteria; 68% were published between 2023 and 2025. Five application domains were identified: (1) clinical documentation and summarisation; (2) clinical decision support and reasoning assistance; (3) patient communication and health-literacy support; (4) biomedical research and knowledge synthesis; and (5) administrative and operational use cases. The evidence base was dominated by benchmark and simulated-workflow studies (89%), with limited prospective workflow-embedded evaluations (11%). Reported benefits concentrated on documentation efficiency, text quality, and knowledge synthesis; safety-relevant risks included hallucinated content, omission of clinically critical information, demographic bias, privacy vulnerabilities, limited explainability, and automation bias. Studies were geographically concentrated in North America and East Asia, with limited representation from Sub-Saharan Africa, South Asia, and Latin America. Current evidence supports cautious deployment of LLMs in selected healthcare tasks under structured oversight. Translational progress depends on prospective evaluation, standardised reporting, equity-focused audits, and lifecycle governance with continuous monitoring. The proposed five-dimensional framework (technical performance, clinical validity, equity, workflow integration, governance) coupled with a three-tier risk model is intended to support researchers and healthcare organisations in assessing readiness and implementing LLM-enabled tools responsibly.

Read PDF

Similar papers

#artificial intelligence Conference Open access Apr 2020

ECCOLA - a Method for Implementing Ethically Aligned AI Systems

The method, ECCOLA, is presented, which aims at making the high-level AI ethics principles more practical, making it possible for developers to more easily implement them in practice.

Ville Vakkuri, Kai-Kristian Kemell, P. Abrahamsson · 64 citations · ⚡6
#computer vision Review Apr 2024

AI-powered Code Review with LLMs: Early Results

The goal is to not only refine the accuracy of the LLM-based tool but also to underscore its potential in streamlining the software development lifecycle through proactive code improvement and education.

Z. Rasheed, Malik Abdul Sami, Muhammad Waseem et al. · 62 citations · ⚡3
#computer vision Open access Mar 2024

LLM-based agents for automating the enhancement of user story quality: An early report

The use of large language models to automatically improve the user story quality in Austrian Post Group IT agile teams is explored, with a reference model for an Autonomous LLM-based Agent System developed and implemented at the company.

Zheying Zhang, M. Rayhan, Tomas Herda et al. · 48 citations · ⚡4
#computer vision Review Mar 2024

System for systematic literature review using multiple AI agents: Concept and an empirical evaluation

This paper introduces a novel multi-AI-agent system designed to fully automate SLRs, and demonstrates how it substantially reduces the time and effort traditionally required for SLRs while maintaining comprehensiveness and precision.

Abdul Malik Sami, Z. Rasheed, Kai-Kristian Kemell et al. · 44 citations · ⚡2
#computer vision Feb 2024

Can Large Language Models Serve as Data Analysts? A Multi-Agent Assisted Approach for Qualitative Data Analysis

The proposed LLM-based multi-agent system automates qualitative data analysis process, creating opportunities for researchers and practitioners, and future improvements focus on enhancing multilingual performance and integrating continuous expert feedback.

Z. Rasheed, Muhammad Waseem, Aakash Ahmad et al. · 41 citations
#artificial intelligence Conference Open access Jun 2018

The Key Concepts of Ethics of Artificial Intelligence

It is suggested that the focus on finding keywords is the first step in guiding and providing direction for future research in the AI ethics field.

Ville Vakkuri, P. Abrahamsson · 39 citations · ⚡2

Related blog posts

Google DeepMind Blog Sep 30, 2026

Introducing SynthID Bio

Proof of concept for watermarking AI-generated proteins while preserving biological function.

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.