Skip to content
Review Open access

Transformer-Based Language Models for Clinical Decision Support Using Clinical Notes: A Scoping Review

Jul 2026 · Information · 0 citations · 49 references

TL;DR

Transformer-based language models have been applied across diverse clinical-note tasks, but the evidence base more strongly supports retrospective task feasibility than transportability, equitable performance, workflow benefit, or safe clinical deployment.

Abstract

Background/Objectives: This scoping review examined recent evidence on the use of transformer-based language models, encompassing encoder-only architectures (e.g., BERT and its clinical variants) and generative large language models (LLMs; e.g., GPT-4 and Llama), to support clinical decision making from unstructured clinical notes, with implications for behavioral-health services where narrative documentation is central. Methods: Following PRISMA-ScR guidelines, PubMed, PsycINFO, and Web of Science were searched for peer-reviewed studies published between 1 January 2023, and 5 August 2025. Studies applying transformer-based language models to clinical narratives for healthcare tasks and reporting evaluative outcomes were included. We extracted data on clinical tasks, model architectures, enhancement strategies, and evaluation metrics; mapped each study by primary purpose, care setting, and primary model approach; and charted reported validation design, direct human comparison, fairness assessment, workflow evaluation, and clinical deployment. Results: Thirty-six studies were included. Information extraction/de-identification and classification/prediction predominated, whereas summarization/generation was less commonly represented. Model approaches appeared to align with task characteristics: encoder-only and decoder-only systems were frequently used for extraction, encoder–decoder systems for generation, and hybrid or pipeline-based approaches for classification and prediction. Standard task-specific metrics (e.g., F1 and AUROC) predominated, whereas evidence beyond retrospective task performance, including direct human comparison, fairness assessment, workflow evaluation, clinical deployment, and temporal or external validation, was rare. No included study evaluated a transformer-based language model application within a behavioral-health service or behavioral-health workflow. Conclusions: Transformer-based language models have been applied across diverse clinical-note tasks, but the evidence base more strongly supports retrospective task feasibility than transportability, equitable performance, workflow benefit, or safe clinical deployment. Future research should prioritize transparent reference standards, external and prospective validation, clinically meaningful human comparison, and equity-focused evaluation, including direct studies in behavioral-health services.

Read PDF

Similar papers

Review Open access Aug 2026

Explainability of decoder-only clinical large language models: A scoping review.

Findings show that clinical LLM explainability has shifted toward fluent generative rationales, but evidence that such explanations reflect model reasoning remains limited, and three regulatory priorities are highlighted: prioritizing explanations that enable independent verification or logic auditing over plausibility-only rationales; preferring inspectable models where regulatory documentation is required; and prospectively validating explanations in clinical workflows before scaling.

Nishant Mishra, Ameen Abu-Hanna, Iacer Calixto · 0 citations
Review Aug 2026

Enhancing Clinical Decision-Making Using Generative AI-Powered Knowledge Retrieval Systems: A Review of Emerging Approaches and Challenges

A wealth of biomedical information and literature, complex electronic health records (EHRs), disjointed guidance, and time pressures associated with the care process are all affecting clinical decision-making. The objective of this review was to discuss the potential of generative AI-driven knowledge retrieval systems for clinical decision-making and to describe some of the technical, compliance, ethical, and implementation challenges and limitations. This purposively selected, 42-source structured narrative review with scoping review elements was conducted based on publications retrieved from PubMed, Scopus, IEEE Xplore, Google Scholar, WHO/FDA, Web of Science, and major clinical informatics journals from 2020-2026. Study results demonstrate that retrieval-augmented generation (RAG) systems can provide accurate and relevant information by grounding content generated by large language models (LLMs) in clinical guidelines, biomedical literature, drug databases, EHR data, and institutional protocols. The results showed that RAG systems yielded better biomedical system performance than the baseline LLM, with an OR of 1.35 (95% CI: 1.19, 1.53). There are potential impacts, however, such as hallucinations, incomplete retrieval, incomplete and comprehensive datasets, privacy breaches, and lack of multilingual validation. The study concludes that evidence-based, auditable, locally adaptable, and supervised by licensed clinician retrieval systems with generative AI can support safer, faster, and more relevant decision-making processes in clinical settings. Future studies should involve prospective multi-site implementation, a clear retrieval pipeline, multilingual datasets, EHR integration, ongoing monitoring, and clinical governance models to ensure safe use.

Sonam Kumari · 0 citations
Review Open access Aug 2026

Large language models in spine care and research.

BACKGROUND Large language models (LLMs) have emerged as powerful transformer-based systems capable of capturing long-range dependencies and complex semantic relationships in clinical language. In this review paper, we first examine the technical foundations of medical LLMs, including transformer architecture, attention mechanisms, training paradigms, and retrieval-augmented generation. RESULTS We then survey documented applications in spine surgery and spinal care, highlighting moderate guideline concordance (46-67%) for diagnostic support, automated generation of operative notes and discharge summaries for administrative workflows, LLM-assisted literature review and manuscript drafting for research support (with ~68% novelty accuracy), and translation of complex surgical concepts into patient-friendly materials at a seventh-grade reading level. We next explore emerging multimodal models that integrate text, imaging, laboratory, and genomic data via cross-modal attention, demonstrating superior performance in holistic diagnostic and prognostic tasks. DISCUSSION Finally, we discuss key implementation challenges, including model accuracy and hallucinations; computational, privacy, and regulatory constraints under HIPAA/GDPR; and bias mitigation, to outline strategies for safe, effective, and equitable deployment. CONCLUSION By mapping technical capabilities to clinical and research use cases, this review highlights the promise of LLMs in enhancing decision support, workflow efficiency, research productivity, and patient communication in spine care, while emphasizing the need for interdisciplinary collaboration, robust evaluation metrics, and governance frameworks that prioritize patient safety and equity.

Fabio Galbusera, Andrea Cina · 0 citations
Review Jul 2026

Large language models in clinical and healthcare scenarios: a global informatics analysis

This paper conducts a comprehensive analysis of evaluation methods, deployment processes, and governance strategies for LLMs in the healthcare field, focusing on three key issues: model version drift, multilingual external validation, and prompt injection security governance.

Song-Bin Guo, Sui-Xing Zhong, Yixian Ma et al. · 0 citations
Review Open access Jul 2026

A Modular Evaluation of AI-Assisted Clinical Documentation

Clinical documentation in Electronic Health Records (EHRs) remains a substantial source of administrative burden for clinicians. In this study, we evaluate a modular AI-assisted clinical documentation pipeline using two complementary approaches: (1) a controlled benchmark based on multilingual synthetic clinical dialogues, and (2) an observational analysis of real-world usage traces from routine deployments. The benchmark enables systematic comparison of ASR–LLM configurations under fully controlled conditions, using metrics for transcription accuracy (Word Error Rate and Medical WER), report-generation quality, and modeled processing cost. Within this benchmark setting, Voxtral showed the strongest ASR performance among the evaluated models, while GPT-4o and Gemini 1.5 Pro showed the strongest report-generation performance under the automated evaluation used in this study. The real-world trace analysis should be interpreted as descriptive evidence of operational use, not as prospective clinical validation or as a direct evaluation of any single benchmarked configuration. Taken together, the results support the use of this pipeline as a human-supervised draft-generation tool that still requires clinician review, local workflow evaluation, and prospective clinical validation before broader deployment.

Julien Delaunay, Maissaa Sarkis, Jordi Solé-Casals et al. · 0 citations