Skip to content
Review Open access

Benchmarking and developing large language models using one million clinical trials

Jul 2026 · npj Digital Medicine · 1 citation · ⚡ 1 influential

TL;DR

A large-scale structured resource aggregating 1.6M clinical trial records from fifteen global registries linked with biomedical ontologies and literature is introduced, demonstrating the potential of domain-adapted AI to improve evidence synthesis and clinical trial design and establishing as a foundation for scaling AI in clinical research.

Abstract

Developing artificial intelligence (AI) for clinical research requires a comprehensive data foundation for model benchmarking and development. Here, we introduce , a large-scale structured resource aggregating 1.6M clinical trial records from fifteen global registries linked with biomedical ontologies and literature. Using this resource, we construct 152K training and testing samples spanning eight clinical research tasks, including systematic review, trial design, and trial optimization. Benchmarking cutting-edge large language models (LLMs) reveals limited clinical reasoning capability in generic LLMs. In contrast, an 8B LLM developed on using supervised fine-tuning and reinforcement learning outperforms 70B generic counterparts across all eight tasks, with relative improvements of 73.7, 67.6, 38.4, 37.8, 26.5, 20.7, 20.0, 18.1, and 5.2%, respectively. These results demonstrate the potential of domain-adapted AI to improve evidence synthesis and clinical trial design, establishing as a foundation for scaling AI in clinical research.

Read PDF

Similar papers

Review Open access Jul 2026

Tutorial: guidance on the use of large language models for medical research

This entry-level tutorial aims to equip healthcare professionals with the tools necessary to effectively integrate LLMs into clinical practice, ensuring that these powerful technologies are applied in a safe, reliable, and impactful manner.

Qiao Jin, Nicholas Wan, Robert Leaman et al. · 1 citation
Review Open access Aug 2026

Large Language Models and Medical AI Systems for Healthcare Diagnosis: A Systematic Review

The rapid growth of artificial intelligence systems (AI systems) has increased interest in the use of patient care and clinical decision-making processes. There is some uncertainty regarding their reliability and safety in clinical practice. A more detailed systematic review of literature examining LLMs applied to healthcare diagnosis was conducted. A PRISMA-based systematic review has been carried out of relevant literature published in the major databases for the years 2022–2025. Key findings include a growing trend to develop multimodal models based on diverse input modalities, combining LLM models with other models as part of clinical workflows. The Usage of complementary methodologies such as retrieval-augmented generation, knowledge graphs, and federated learning is highly expanding, particularly in enhancing the efficiency and accuracy of clinical decision-making processes. Significant challenges such as hallucinations, bias, prompt sensitivity, limited explainability, and inadequate clinical validation continue to pose major obstacles. Although promising, LLM-based systems are not yet reliable enough for autonomous medical diagnosis. Overall, this review contains multiple recommendations for future research in many areas (e.g., LLMs) to ensure a high level of safety, transparency, and clinical applicability for LLMs and other AI/ML-related technologies and devices.

M. U. K. Gunawardhna, Pirunthavi Wijikumar, D. Weerasinghe · 0 citations
Preprint Aug 2026

A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower, and this results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves grounding at some cost to communication polish.

Praveen Reddy, C. Mandke, Suvrankar Datta et al. · 0 citations
Review Open access Jul 2026

Artificial intelligence for genomic science: a scoping review of concepts, architectures, applications, and open challenges

Introduction Artificial intelligence (AI) is becoming central to genomics and multi-omics, but its concepts, architectures, applications, evaluation standards, and translational requirements remain fragmented. This scoping review mapped how AI is defined and operationalized in genomic science, including machine learning, deep learning, graph-based methods, foundation models, and large language models, and synthesized their data modalities, applications, evaluation practices, interpretability strategies, and governance challenges. Methods We conducted a PRISMA-ScR scoping review with Joanna Briggs Institute guidance. Eligible studies applied AI to genomics or closely allied omics in research, clinical, or public health contexts. MEDLINE/PubMed, Embase, and supplementary registers were searched from January 2001 to 3 September 2025 without language restrictions. Records were screened in duplicate, and standardized items were extracted, including AI concept or method family, omics modality, task, metrics, interpretability, governance, and deployment considerations. Methodological reporting and quality were appraised using design-appropriate JBI tools and summarized descriptively as a normalized 0%–100% checklist-fulfillment index. Results From 3,785 records, 1,040 studies were included. Publication remained sparse until 2017 and then expanded steeply, with more than 90% appearing from 2018 onward. The normalized JBI checklist-fulfillment index was modest overall (mean 35.3%, SD 20.1; range 7.5%–87.5%) and was interpreted descriptively, not as a directly comparable quality score across designs. Conceptually, the field has moved from feature-engineered statistical learning toward representation learning systems modeling nucleotide sequences, regulatory context, single-cell states, multi-omics profiles, biomedical text, and clinical-genomic knowledge. Applications concentrated on variant interpretation, regulatory genomics, multi-omics integration, single-cell analysis, pathology/radiology-genomics fusion, and genomic decision support, with increasing use of deep learning, graph models, foundation models, and LLMs. Calibration, external validation, mechanistic interpretability, ancestry-aware fairness, privacy protection, and deployment models for sensitive genomic data were unevenly reported; prospective multisite evaluations were rare. Discussion AI in genomics has scaled rapidly since 2017–2018, but translation remains constrained by heterogeneous concepts, inconsistent benchmarks, incomplete reporting, and limited governance. Priorities include biologically meaningful benchmarks; calibrated uncertainty for genomic decision support; mechanism-linked interpretability; ancestry- and site-aware validation; privacy-preserving analysis of sensitive genomic data; and human oversight for variant interpretation, precision medicine, and public health genomics. Systematic Review Registration https://osf.io/uexzh.

W. Rodrigues, M. Parise, Doglas Parise et al. · 1 citation
Review Jul 2026

Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation

Large language models and multimodal foundation models are enabling medical artificial intelligence (AI) systems to move beyond isolated prediction and undertake multistep clinical tasks that require planning, tool use, memory, iterative correction, and coordination among specialized agents. However, the scope of agentic AI in medicine remains unsettled, and current evaluation practices are not yet aligned with the requirements of clinical use. We conducted a scoping review with systematic evidence mapping across five electronic sources, screened 1,649 exportable records, and provisionally included 557 unique studies that met predefined criteria for goal-directed task execution, tool use, interaction with external resources, feedback-based refinement, or multi-agent collaboration. The included studies describe single agents that use external tools, workflows supported by retrieval and external knowledge, multimodal agents, and multi-agent systems applied to medical question answering, image interpretation, electronic health record analysis, drug safety, and clinical trial prediction. The evidence base remains dominated by public benchmarks, simulated settings, retrospective datasets, and small-scale expert evaluation. Process reliability, evidence traceability, uncertainty, safety, workflow impact, and external validity are evaluated less consistently. Clinical translation will depend on clearer definitions, reproducible evaluation, auditable oversight, interoperable system design, and prospective validation in real-world clinical workflows.

Zheng Tong, Yang Liu, Wanshu Fan et al. · 0 citations