Skip to content
Review Open access

Artificial Intelligence and Generative Models in Hepatology: From Large Language Models to Digital Pathology in Liver Disease Diagnosis and Treatment.

Jul 2026 · Clinical and Molecular Hepatology · 0 citations · 93 references
Medicine

TL;DR

This narrative review examines the transition from task-specific discrimination AI to large language models (LLMs), multimodal foundation models, and agentic AI, and outlines the future directions for safe AI model deployment based on multimodal foundation models, prospective and federated evaluation, lifecycle governance, and continuous monitoring for performance, calibration, and equity.

Abstract

Artificial intelligence (AI), particularly foundation and generative models, is reshaping the practice of hepatology through enhanced knowledge synthesis, quantitative and reproducible analysis of multimodal data, and personalized clinical decision support. This narrative review examines the transition from task-specific discrimination AI to large language models (LLMs), multimodal foundation models, and agentic AI. We synthesize evidence from original and validation studies, clinical evaluations, and benchmark studies, as well as expert reviews and regulatory frameworks across metabolic dysfunction-associated steatotic liver disease, chronic hepatitis B, cirrhosis and portal hypertension, hepatocellular carcinoma, and liver transplantation. LLMs can convert free-text notes into structured data, summarize longitudinal electronic health records, support patient education, and retrieve guideline-based information. Retrieval-augmented generation and agentic AI may improve traceability and workflow support, but current evidence is largely retrospective or proof-of-concept. In digital pathology and imaging, discriminative AI has enabled more quantitative and reproducible histologic scoring and biomarker analysis. Pathology and multimodal foundation models offer transferable representations, report generation, and cross-modal reasoning, but hepatology-specific validation remains limited. Key risks include hallucination, automation bias, domain shift across centers and devices, and inequities due to under-representation of patient subgroups. We outline the future directions for safe AI model deployment based on multimodal foundation models, prospective and federated evaluation, lifecycle governance, and continuous monitoring for performance, calibration, and equity. Most generative AI applications in hepatology remain at the proof-of-concept stage, and rigorous prospective validation with human-in-the-loop oversight is required before clinical integration.

Read PDF

Similar papers

Review Open access Sep 2026

Generative artificial intelligence in lung cancer care: current applications, challenges, and future directions

Generative artificial intelligence (GAI), particularly large language models (LLMs) and multimodal foundation models, represents a new generation of artificial intelligence technologies with emerging applications in healthcare. In lung cancer care, where clinical decisions increasingly require integration of imaging, pathology, molecular alterations, and rapidly evolving therapeutic evidence, GAI provides new opportunities to enhance clinical information synthesis and decision support. Recent studies have explored the application of GAI across the lung cancer care continuum, including screening and early detection, diagnosis and characterization, treatment decision-making, prognosis and disease monitoring, patient–clinician communication, and clinical workflow optimization. For example, multimodal foundation models such as M3FM have demonstrated the feasibility of jointly learning imaging and clinical information to support multiple lung cancer-related tasks within a unified architecture. In addition, oncology-specific LLMs trained on real-world clinical data have shown promise in predicting lung cancer progression by integrating longitudinal radiological and clinical information. However, most current applications remain at an exploratory or early validation stage, with substantial heterogeneity in model architectures, evaluation frameworks, and levels of clinical validation. Challenges related to evidence quality, data integration, safety, transparency, and regulatory oversight must be addressed before widespread clinical implementation. In this review, we provide a clinically oriented overview of the current applications of GAI across the lung cancer care continuum and discuss emerging developments, limitations, and future directions, including domain-specific models, multimodal systems, guideline-integrated decision support, and prospective validation frameworks.

Wen-Zheng Zhang, Zhao-Rui Feng, Yi-Tong Liu et al. · 0 citations
#federated learning Review Open access Sep 2026

Applications and challenges of artificial intelligence in cancer management

Artificial intelligence (AI) is increasingly transforming cancer management by enabling the analysis of complex multimodal data generated across the cancer care continuum. This structured narrative review synthesizes current applications of AI, machine learning, and deep learning in cancer detection, imaging diagnosis, tumor characterization, histopathology, genomics, drug discovery, precision oncology, prognosis, treatment-response prediction, clinical decision support, adherence monitoring, and supportive care. A literature search was conducted using PubMed, Scopus, Web of Science, IEEE Xplore, and Google Scholar for publications from July 2001 to November 2025. From 564 identified records, 234 articles were included in the final thematic synthesis after duplicate removal, title/abstract screening, and full-text assessment. The reviewed evidence shows that AI models, including convolutional neural networks, U-Net-based segmentation models, radiomics approaches, hybrid models, multiple-instance learning, federated learning, large language models, and multimodal frameworks, have demonstrated substantial potential in extracting clinically meaningful patterns from imaging, pathology, genomic, biomarker, and clinical data. These tools may support earlier detection, improved tumor classification, non-invasive molecular profiling, treatment optimization, and personalized decision-making. However, clinical readiness remains uneven, as many models are limited by retrospective designs, dataset heterogeneity, data leakage, overfitting, limited explainability, insufficient external validation, and uncertain generalizability across global healthcare settings. Future progress will require diverse multicenter datasets, standardized annotation, transparent reporting, explainable and hybrid modelling, ethical data governance, regulatory oversight, and post-deployment monitoring. AI should therefore be viewed as clinician-supervised decision support that can strengthen precision oncology when implemented responsibly.

Eva Rahman Kabir, N. Mustafa, Zara Sheikh et al. · 0 citations
Review Open access Apr 2026

Large language models in hepatology: A systematic review

The authors' analysis reveals that LLMs demonstrate promising capabilities in processing textual and visual data related to various liver diseases, including hepatocellular carcinoma, cirrhosis, and non-alcoholic fatty liver disease, but study heterogeneity and significant challenges remain regarding accuracy, reliability, and safety.

T. Suenghataiphorn, Narisara Tribuddharat, Pojsakorn Danpanichkul et al. · 0 citations
Review Open access Jul 2026

Large AI Models in Life Sciences and Healthcare: A Review and Analysis

Large AI models are reshaping biomedical research and healthcare delivery by promoting more intelligent, data-driven, and integrated approaches by demonstrating superior capabilities in knowledge representation, reasoning, and cross-domain learning.

Boyuan Fang · 0 citations
Review Open access Aug 2026

Explainable artificial intelligence in prostate and bladder cancer: A review of imaging, digital pathology, molecular profiling, and clinical data

Artificial intelligence (AI) models are increasingly applied across the detection, staging, treatment selection, and surveillance of prostate and bladder cancer. Explainable AI links predictions to identifiable features, spatial regions, or architectural constraints, supporting clinical audit and informed use. This narrative review synthesizes evidence across both cancers in imaging, digital pathology, molecular and liquid-biopsy data, and clinical applications and assesses deployment readiness. Two parallel structured searches of PubMed/MEDLINE were performed, one for each cancer, covering records through November 15, 2025. Eligible studies applied machine or deep learning to a clinical or translational task and included an explainability component. Seventy studies across 4 data domains were included. Imaging was the most active domain. SHapley Additive exPlanations and gradient-weighted class activation mapping confirmed biologically plausible signal localization. Interpretability-by-design in bladder endoscopy embedded explainability without post hoc approximation. Attention heatmaps and probability maps in digital pathology linked model focus to histological features relevant to grading and invasion. Molecular models ranged from post hoc attribution to pathway-constrained architectures; a biologically structured deep neural network in prostate cancer demonstrated how architectural constraints can move explanations toward mechanistic inference. Clinical and electronic medical record-based models constituted the largest group, using SHapley Additive exPlanations for prebiopsy triage, treatment selection, recurrence prediction, and active surveillance. Post hoc explanations and interpretable-by-design models carry different evidentiary weight. Prospective evidence that explainable AI affects clinical decisions is lacking. Translation requires independent validation of fidelity, stability, and usability, prospective reader studies, standardized reporting, and investment in interpretable architectures.

D. Diamantidis, Georgios Tsakaldimis, Nikolaos Smyrlis et al. · 0 citations
Open access Sep 2026

Large Language Models for Heterogeneous Data Mining in Liver Disease: Framework Development and Retrospective Validation Study

Abstract Background Differentiating among liver disease entities such as autoimmune liver disease (AILD), drug-induced liver injury (DILI), and chronic hepatitis B (CHB) remains clinically challenging due to overlapping clinical manifestations and nonspecific laboratory findings. Conventional machine learning (ML) approaches rely mainly on structured laboratory data, whereas free-text clinical reports and other heterogeneous electronic medical record data are often underused. Large language models (LLMs) may provide a strategy for encoding heterogeneous clinical information, yet their usefulness for liver disease classification remains insufficiently evaluated. Objective This study aimed to evaluate the usefulness of LLM-derived embeddings for clinical data mining in liver disease and to determine whether integrating these embeddings with laboratory variables improves classification across broad disease categories and closely related subtypes. Methods We retrospectively analyzed electronic medical record data from 7543 patients with nonoverlapping liver disease etiologies treated at Beijing Youan Hospital, Capital Medical University, between 2010 and 2025. Three LLMs (Qwen3, Huatuo-o1, and II-Medical) generated semantic embeddings from standardized clinical text, combining free-text examination reports, and structured clinical observations. Performance was assessed in a 3-class etiological task (AILD, DILI, and CHB) and a 4-class task further subclassifying AILD into autoimmune hepatitis and primary biliary cholangitis. We compared embedding-only models, LLM-integrated ML models, and an ML-only baseline using the same structured variable set and preprocessing pipeline, with lightweight natural language processing encoders and zero-shot LLM reasoning as additional comparators. Models were developed using 5-fold cross-validation and evaluated on an internal holdout set using accuracy, macroaveraged precision, recall, and F1-score. Results In the 3-class task, the LLM-integrated ML models achieved macro F1-scores of 0.835‐0.837, compared with 0.791 for the ML-only baseline, with corresponding accuracies of 0.925‐0.929 versus 0.893. In the 4-class task, the LLM-integrated ML models achieved macro F1-scores of 0.717‐0.734, compared with 0.665 for the ML-only baseline, with corresponding accuracies of 0.920‐0.922 versus 0.874. A temporal split sensitivity analysis using cases from 2010 to 2019 for training and cases from 2020 to 2025 for testing showed that the relative advantage of LLM-integrated ML models over the ML-only baseline was preserved. Direct zero-shot LLM reasoning and lightweight natural language processing encoders performed below the embedding-based integrated models. Conclusions In this single-center retrospective cohort of patients with clear-cut, nonoverlapping liver disease etiologies, LLM-derived embeddings provided complementary information to structured laboratory variables for multiclass liver disease classification. The integrated framework showed improved internal validation performance compared with the ML-only model, particularly for non-CHB categories and fine-grained subtype discrimination. Because patients with overlapping liver disease etiologies were excluded, the reported performance may overestimate diagnostic accuracy in broader real-world clinical settings where overlapping syndromes are common. Multicenter external validation and prospective evaluation in more heterogeneous patient populations are needed before clinical implementation.

Hai-Ping Zhang, Xin-Ming Li, Ke-Chi Fang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.