Transformer-based language models have been applied across diverse clinical-note tasks, but the evidence base more strongly supports retrospective task feasibility than transportability, equitable performance, workflow benefit, or safe clinical deployment.
Abstract
Background/Objectives: This scoping review examined recent evidence on the use of transformer-based language models, encompassing encoder-only architectures (e.g., BERT and its clinical variants) and generative large language models (LLMs; e.g., GPT-4 and Llama), to support clinical decision making from unstructured clinical notes, with implications for behavioral-health services where narrative documentation is central. Methods: Following PRISMA-ScR guidelines, PubMed, PsycINFO, and Web of Science were searched for peer-reviewed studies published between 1 January 2023, and 5 August 2025. Studies applying transformer-based language models to clinical narratives for healthcare tasks and reporting evaluative outcomes were included. We extracted data on clinical tasks, model architectures, enhancement strategies, and evaluation metrics; mapped each study by primary purpose, care setting, and primary model approach; and charted reported validation design, direct human comparison, fairness assessment, workflow evaluation, and clinical deployment. Results: Thirty-six studies were included. Information extraction/de-identification and classification/prediction predominated, whereas summarization/generation was less commonly represented. Model approaches appeared to align with task characteristics: encoder-only and decoder-only systems were frequently used for extraction, encoder–decoder systems for generation, and hybrid or pipeline-based approaches for classification and prediction. Standard task-specific metrics (e.g., F1 and AUROC) predominated, whereas evidence beyond retrospective task performance, including direct human comparison, fairness assessment, workflow evaluation, clinical deployment, and temporal or external validation, was rare. No included study evaluated a transformer-based language model application within a behavioral-health service or behavioral-health workflow. Conclusions: Transformer-based language models have been applied across diverse clinical-note tasks, but the evidence base more strongly supports retrospective task feasibility than transportability, equitable performance, workflow benefit, or safe clinical deployment. Future research should prioritize transparent reference standards, external and prospective validation, clinically meaningful human comparison, and equity-focused evaluation, including direct studies in behavioral-health services.
Findings show that clinical LLM explainability has shifted toward fluent generative rationales, but evidence that such explanations reflect model reasoning remains limited, and three regulatory priorities are highlighted: prioritizing explanations that enable independent verification or logic auditing over plausibility-only rationales; preferring inspectable models where regulatory documentation is required; and prospectively validating explanations in clinical workflows before scaling.
A wealth of biomedical information and literature, complex electronic health records (EHRs), disjointed guidance, and time pressures associated with the care process are all affecting clinical decision-making. The objective of this review was to discuss the potential of generative AI-driven knowledge retrieval systems for clinical decision-making and to describe some of the technical, compliance, ethical, and implementation challenges and limitations. This purposively selected, 42-source structured narrative review with scoping review elements was conducted based on publications retrieved from PubMed, Scopus, IEEE Xplore, Google Scholar, WHO/FDA, Web of Science, and major clinical informatics journals from 2020-2026. Study results demonstrate that retrieval-augmented generation (RAG) systems can provide accurate and relevant information by grounding content generated by large language models (LLMs) in clinical guidelines, biomedical literature, drug databases, EHR data, and institutional protocols. The results showed that RAG systems yielded better biomedical system performance than the baseline LLM, with an OR of 1.35 (95% CI: 1.19, 1.53). There are potential impacts, however, such as hallucinations, incomplete retrieval, incomplete and comprehensive datasets, privacy breaches, and lack of multilingual validation. The study concludes that evidence-based, auditable, locally adaptable, and supervised by licensed clinician retrieval systems with generative AI can support safer, faster, and more relevant decision-making processes in clinical settings. Future studies should involve prospective multi-site implementation, a clear retrieval pipeline, multilingual datasets, EHR integration, ongoing monitoring, and clinical governance models to ensure safe use.
Sonam Kumari· International Journal of Adv...· 0 citations
BACKGROUND
Large language models (LLMs) have emerged as powerful transformer-based systems capable of capturing long-range dependencies and complex semantic relationships in clinical language. In this review paper, we first examine the technical foundations of medical LLMs, including transformer architecture, attention mechanisms, training paradigms, and retrieval-augmented generation.
RESULTS
We then survey documented applications in spine surgery and spinal care, highlighting moderate guideline concordance (46-67%) for diagnostic support, automated generation of operative notes and discharge summaries for administrative workflows, LLM-assisted literature review and manuscript drafting for research support (with ~68% novelty accuracy), and translation of complex surgical concepts into patient-friendly materials at a seventh-grade reading level. We next explore emerging multimodal models that integrate text, imaging, laboratory, and genomic data via cross-modal attention, demonstrating superior performance in holistic diagnostic and prognostic tasks.
DISCUSSION
Finally, we discuss key implementation challenges, including model accuracy and hallucinations; computational, privacy, and regulatory constraints under HIPAA/GDPR; and bias mitigation, to outline strategies for safe, effective, and equitable deployment.
CONCLUSION
By mapping technical capabilities to clinical and research use cases, this review highlights the promise of LLMs in enhancing decision support, workflow efficiency, research productivity, and patient communication in spine care, while emphasizing the need for interdisciplinary collaboration, robust evaluation metrics, and governance frameworks that prioritize patient safety and equity.
Fabio Galbusera, Andrea Cina· European spine journal· 0 citations
This paper conducts a comprehensive analysis of evaluation methods, deployment processes, and governance strategies for LLMs in the healthcare field, focusing on three key issues: model version drift, multilingual external validation, and prompt injection security governance.
Song-Bin Guo, Sui-Xing Zhong, Yixian Ma et al.· International Journal of Sur...· 0 citations
A scoping review of 24 PubMed-indexed studies published between 2023 and 2026 was conducted to assess current applications, benefits, limitations, and future directions of LLMs in healthcare.
Antoni Klamka, Paulina Kawalec, Kamil Bronikowski et al.· Quality in Sport· 0 citations
Clinical documentation in Electronic Health Records (EHRs) remains a substantial source of administrative burden for clinicians. In this study, we evaluate a modular AI-assisted clinical documentation pipeline using two complementary approaches: (1) a controlled benchmark based on multilingual synthetic clinical dialogues, and (2) an observational analysis of real-world usage traces from routine deployments. The benchmark enables systematic comparison of ASR–LLM configurations under fully controlled conditions, using metrics for transcription accuracy (Word Error Rate and Medical WER), report-generation quality, and modeled processing cost. Within this benchmark setting, Voxtral showed the strongest ASR performance among the evaluated models, while GPT-4o and Gemini 1.5 Pro showed the strongest report-generation performance under the automated evaluation used in this study. The real-world trace analysis should be interpreted as descriptive evidence of operational use, not as prospective clinical validation or as a direct evaluation of any single benchmarked configuration. Taken together, the results support the use of this pipeline as a human-supervised draft-generation tool that still requires clinician review, local workflow evaluation, and prospective clinical validation before broader deployment.