While AI agents delivered highly efficient, directionally aligned assessments, they did not fully capture the nuances of human clinical judgment and could not substitute for physician-centered evaluation and promise assistive tools that can triage or pre-screen outputs to reduce human burden.
Abstract
While multimodal large language models (LLMs) demonstrate significant potential in healthcare applications, their clinical utility is difficult to appraise. Current evaluations of medical-assisting LLMs are often limited by sparse human expertise, narrow specialty scope, and reliance on multiple-choice benchmarks or synthetic vignettes, which can inflate performance and obscure clinical utility. We conducted a multicenter, multidisciplinary study in which more than 400 physicians-spanning seven specialties, varied experience levels, and multiple geographic settings-evaluated LLM-generated free-text responses to real, de-identified clinical cases. In a matched-control design, we also deployed an equivalent number of AI agents configured to mirror physician characteristics to examine whether automated evaluators can supplement or replace human assessment. Our results demonstrated that physician assessments exhibited substantial heterogeneity by clinical seniority and practice environment, leading to notable shifts in relative model rankings across cohorts. While AI agents delivered highly efficient, directionally aligned assessments, they did not fully capture the nuances of human clinical judgment and could not substitute for physician-centered evaluation. Instead, they promise assistive tools that can triage or pre-screen outputs to reduce human burden.
The application of Large Language Models (LLMs) to diagnostic decision-making has garnered growing interest. However, existing benchmarks largely focus on textual reasoning or isolated visual question-answering (VQA) tasks, lacking holistic integration of clinical narratives and medical imaging, and thus failing to assess the multimodal diagnostic synthesis capability central to expert clinical judgment. To bridge this gap, we introduce MedReaMM, a benchmark specifically designed to evaluate models'ability to synthesize heterogeneous clinical evidence consisting of detailed patient histories alongside multiple medical images into accurate differential diagnoses under a complete-information paradigm. Constructed from case reports sourced from top-tier medical journals and curated clinical case databases, MedReaMM comprises 625 expert-validated cases with an average of 2.79 medical images per case and a total of 1,042 standardized diagnoses annotated with ICD-11 codes. These cases predominantly represent rare, atypical, or multi-system presentations that demand expert-level evidence integration beyond routine pattern recognition. We evaluate 23 Large Multimodal Models (LMMs) and find that most achieve diagnostic accuracy scores below 50%, underscoring a substantial gap in multimodal diagnostic synthesis capability. Further analysis reveals that medical knowledge proficiency, medical image understanding, and evidence integration are all highly correlated with diagnostic performance.
Lai Wei, Yuchao Chen, Zhenbiao Cao et al.· 0 citations
The rapid growth of artificial intelligence systems (AI systems) has increased interest in the use of patient care and clinical decision-making processes. There is some uncertainty regarding their reliability and safety in clinical practice. A more detailed systematic review of literature examining LLMs applied to healthcare diagnosis was conducted. A PRISMA-based systematic review has been carried out of relevant literature published in the major databases for the years 2022–2025. Key findings include a growing trend to develop multimodal models based on diverse input modalities, combining LLM models with other models as part of clinical workflows. The Usage of complementary methodologies such as retrieval-augmented generation, knowledge graphs, and federated learning is highly expanding, particularly in enhancing the efficiency and accuracy of clinical decision-making processes. Significant challenges such as hallucinations, bias, prompt sensitivity, limited explainability, and inadequate clinical validation continue to pose major obstacles. Although promising, LLM-based systems are not yet reliable enough for autonomous medical diagnosis. Overall, this review contains multiple recommendations for future research in many areas (e.g., LLMs) to ensure a high level of safety, transparency, and clinical applicability for LLMs and other AI/ML-related technologies and devices.
M. U. K. Gunawardhna, Pirunthavi Wijikumar, D. Weerasinghe· Sri Lankan Journal of Applie...· 0 citations
Errors in differential diagnosis often arise while clinicians are generating and comparing candidate explanations. This review examines the use of large language models (LLMs) for this part of diagnostic reasoning. Internal medicine and pediatrics are the main focus; evidence from radiology, surgical subspecialties, infectious disease, and mental health is used to examine how findings change across specialties. Reported performance depends on the clinical setting, the quality of the input, the prompt, model adaptation, and the evaluation design. Some studies place LLMs near trainees and find that they produce wider, better-organized differentials. Experienced clinicians, however, remain more reliable overall. Domain adaptation, external knowledge, and interactive workflows have improved performance in specific evaluations, but hallucinations and automation bias remain, alongside unresolved questions of governance. Current evidence therefore supports clinician-supervised use of artificial intelligence (AI) systems rather than autonomous diagnosis, pending prospective and specialty-specific evaluation.
Yun-Jia Wu, Qi Yan, Dingcheng Tian· AI Medicine· 0 citations
Uncommon and off-guideline cases are difficult for clinical decision support, because physicians must make a series of management decisions under diagnostic uncertainty and rarely see the full case at once. Most large language model (LLM) benchmarks for medicine score only the final diagnosis, yet much of clinical care turns on the next appropriate action: the next test to order, the imaging study to obtain, the specialist to involve, or the differential to pursue. We introduce MedUPSQA, a dataset of 21,874 mid-stream clinical decision points built from 5,535 real case reports, and MedUPS, an alignment framework that supervises models on these intermediate decisions as they unfold along a patient's trajectory. We segment free-text case presentations into chronologically ordered, accumulating clinical chunks and align models to predict the next step with reinforcement learning (GRPO), using an external LLM-as-a-Judge reward. This objective mirrors how clinicians actually meet patients, reasoning forward from accumulating evidence toward the next decision, rather than committing to a final label. Across three backbones, mid-stream alignment raises next-step accuracy from 55.2 to 66.7 for Qwen3.6-27B, from 47.2 to 57.8 for Qwen3.5-9B, and from 37.8 to 44.4 for HuatuoGPT-3-8B, with 95% CI. In several model scales we test the objective improves accuracy more than scale, with smaller models surpassing larger, frontier models we evaluate. We further train supervised fine-tuning (SFT) baselines on the mid-stream task, SFT improves all backbones above base, indicating the target framwork carries signal independently of the optimizer. We release the dataset, code, and aligned checkpoints.
Ofir Ben Shoham, O. Perets, Nir Grinberg et al.· 0 citations
Artificial intelligence (AI) predictive models demonstrate potential for transforming clinical decision-making across medicine. However, conventional randomized controlled trials (RCTs), the gold standard for evaluating medical interventions, are ill-suited for clinical AI tools due to their static design, lengthy timelines, and inability to accommodate algorithms that evolve and adapt to changing clinical contexts. In this narrative review, we outline the limitations of traditional evaluation frameworks and propose a paradigm shift toward adaptive, iterative, and context-specific assessment methodologies. Evaluating clinical AI in practice requires three interdependent but epistemologically distinct activities: performance monitoring, which tracks the technical characteristics of the deployed model (calibration, discrimination, data drift, alert burden, fairness, workflow fidelity); clinical impact monitoring, which observationally and prospectively tracks whether the initial clinical benefit appears sustained over time; and scientific evidence generation, which produces causal estimates of deployment effects on patient outcomes through pragmatic, adaptive trial designs and causal inference techniques. We propose a predictive-AI-specific framework that links performance monitoring, clinical impact monitoring, evidence generation, causal estimands, and governance of model updates into one coherent decision pathway for clinicians and trialists. We present a governance-driven escalation protocol specifying when monitoring signals should trigger formal evidence generation, a decision pathway mapping signal types (performance or clinical impact) to trial design classifications, and a guide to causal inference methods for clinical AI trials. Drawing from adaptive platform and pragmatic trial designs, we recommend continuous monitoring approaches that prioritize patient-centered outcomes, health equity, and workflow integration over narrow performance metrics, and provide actionable steps to design a clinical AI trial. Successful implementation requires clinician engagement, transparency, and ongoing education regarding AI capabilities and limitations. Within this new evaluation paradigm, predictive AI can progress from a promising technology to reliable clinical tools that improve patient outcomes, support clinical decision-making, and uphold ethical standards in routine practice.
M. Fosset, Joris Pensier, Boris Jung et al.· PLOS Digital Health· 0 citations
Clinical reasoning and differential diagnosis are core competencies in medicine. Large language models (LLMs) have generated considerable interest as potential tools to support these skills. This article presents a narrative review of the available evidence, organized around five key questions: the effect of LLMs on diagnostic reasoning, the optimal design of clinician-LLM interaction, the appropriate timing of consultation during the clinical encounter, the safest models of clinical-AI integration, and the main risks associated with their use. The evidence shows that LLMs improve differential diagnosis when used by trained professionals within structured workflows. However, passive use generates biases, and clinician-AI collaboration may not consistently outperform autonomous LLMs. A practical framework stratified by degree of diagnostic uncertainty is proposed, with operational and educational recommendations oriented toward "physician-in-the-loop" models, in which LLMs amplify, challenge, and make explicit the diagnostic reasoning process under critical human oversight.
L. Corral-Gudino, M. Ramos-Casals, Miguel Marcos et al.· Medicina clínica (Ed. impres...· 0 citations