Aug 2026· Indian Journal of Computer Science and Technology· 0 citations
TL;DR
A reproducible estimate of diagnostic retrieval accuracy across four widely used model configurations is provided to establish a baseline for further clinical validation and establish a baseline for further clinical validation.
Abstract
Large language models show potential for clinical diagnostic support, but their diagnostic accuracy across diverse real-world patient presentations remains uncertain. We evaluated diagnostic retrieval and ranking using multi-system emergency-department narratives from MIMIC-IV-Ext version 1.0.2, a deidentified research dataset derived from MIMIC-IV and curated for research involving referral, triage and diagnostic prediction. The dataset was selected because it links early clinical information, including presenting complaints, history and initial vital signs, with documented diagnoses derived from routine care. A locked cohort of 995 diagnosis-free vignettes was used, with one protected index primary diagnosis per case. GPT-5.6 Thinking, Claude Sonnet 5, Llama-3.1-8B-Instruct and Mistral-7B-Instruct-v0.3 independently generated exactly three ranked differential diagnoses for every vignette. The principal outcome was concept-equivalent Top-3 accuracy; Top-1 accuracy, mean reciprocal rank, omission rate, strict text matching, system-wise performance and paired statistical comparisons were secondary outcomes. GPT-5.6 achieved the highest Top-1 accuracy (41.7%), Top-3 accuracy (61.1%) and mean reciprocal rank (0.504). Claude ranked second at 39.8%, 55.3% and 0.467, respectively. Llama reached 25.5% Top-1 and 41.3% Top-3 accuracy, while Mistral reached 22.8% and 38.7%. Overall Top-3 outcomes differed significantly across models (Cochran Q=271.73, df=3, p=1.31×10⁻⁵⁸). Under identical clinical inputs and scoring rules, the proprietary models retrieved the documented index diagnosis more often and ranked it higher than the two open-weight models. These findings provide a reproducible estimate of diagnostic retrieval accuracy across four widely used model configurations and establish a baseline for further clinical validation.
Abstract Background Although large language models (LLMs) have demonstrated the ability to generate the impression section from radiology findings automatically, the incremental diagnostic value of clinical information for these models remains unclear. Objective This study aimed to evaluate the incremental diagnostic value of clinical information for LLMs and compare their performance with that of radiologists. Methods This retrospective study included radiology reports from patients with histopathologically confirmed liver, lung, and breast diseases from 2 institutions between October 2021 and February 2025. We defined three progressive information input scenarios: (1) basic patient information and imaging findings, (2) scenario A plus chief complaint or clinical history, and (3) scenario B plus key laboratory results. Scenario-based data were input into 3 general-purpose LLMs (DeepSeek-R1, Gemini 2.5 Pro, and GPT-4o), generating 2709 entries. Diagnostic accuracy was assessed for both benign-malignant differentiation and disease diagnosis, with histopathology serving as the reference standard. Accuracy was compared among scenarios and against radiologist performance using the McNemar test, and P values were adjusted using the Holm-Bonferroni correction for multiple comparisons. Results A total of 301 patients with pathologically confirmed diseases were included (mean age 53.5, SD 12.0 years; women: n=208, 69.1%). In the liver cohort, a numerical trend toward higher accuracy was observed in scenario C compared with scenario A across all 3 models (scenario C range: 72.3%‐76.2% vs scenario A range: 64.4%‐68.3%); these differences did not reach statistical significance after Holm-Bonferroni correction (all adjusted P>.99). Notably, the DeepSeek-R1 model in scenario C achieved the highest diagnostic accuracy (77/101, 76.2%), with no evidence of a difference compared with radiologists (82/101, 81.2%; adjusted P>.99). In contrast, results in the lung and breast cohorts were more heterogeneous. In the lung cohort, GPT-4o achieved its highest accuracy in scenario A for disease diagnosis (68/92, 73.9%), which exceeded its performance in scenario B (64/92, 69.6%) and scenario C (66/92, 71.7%), suggesting that additional clinical information did not confer a consistent benefit. Gemini 2.5 Pro in scenario B achieved the highest accuracy in this cohort (72/92, 78.3%); however, no statistically significant difference was found compared with radiologists (80/92, 87.0%; adjusted P=.25). In the breast cohort, DeepSeek-R1 achieved the numerically highest diagnostic accuracy in scenario A, and it decreased numerically with the addition of laboratory tests, although no significant difference was found between scenarios A and C (73/108, 67.6% vs 71/108, 65.7%; adjusted P>.99). Conclusions While the addition of clinical information was associated with a numeric trend toward higher diagnostic accuracy overall, this trend was heterogeneous across models and disease types, and no statistically significant improvement was demonstrated after adjustment for multiple comparisons.
Jin-Qi Zhang, Xiao-Yi Wang, Yanfeng Zhao et al.· Journal of Medical Internet...· 0 citations
Errors in differential diagnosis often arise while clinicians are generating and comparing candidate explanations. This review examines the use of large language models (LLMs) for this part of diagnostic reasoning. Internal medicine and pediatrics are the main focus; evidence from radiology, surgical subspecialties, infectious disease, and mental health is used to examine how findings change across specialties. Reported performance depends on the clinical setting, the quality of the input, the prompt, model adaptation, and the evaluation design. Some studies place LLMs near trainees and find that they produce wider, better-organized differentials. Experienced clinicians, however, remain more reliable overall. Domain adaptation, external knowledge, and interactive workflows have improved performance in specific evaluations, but hallucinations and automation bias remain, alongside unresolved questions of governance. Current evidence therefore supports clinician-supervised use of artificial intelligence (AI) systems rather than autonomous diagnosis, pending prospective and specialty-specific evaluation.
Yun-Jia Wu, Qi Yan, Dingcheng Tian· AI Medicine· 0 citations
Background: Large language models (LLMs) have shown increasing capability in medical knowledge tasks, yet how they perform in extracting structured clinical information from real-world clinical documentation remains uncertain. We evaluated the performance of LLMs relative to medical professionals in extracting SNOMED-coded clinical information from openly available Ear, Nose and Throat (ENT) EHRs from MTSamples, examining both reliability and accuracy metrics. Methods: We evaluated the performance of seven LLMs (including GPT-4o, Claude 3.5, Gemini 1.5 Pro, Gemma 3 and three LLAMA variants) against annotations from fourteen medical professionals who served as both study authors and data annotators. Each annotator independently extracted seven categories of clinical information from 98 publicly available ENT clinical documents: socio-demographics, symptoms, signs, diagnoses, treatments, risk factors, and test results. Standardised medical terminology was enforced through SNOMED-CT code assignment, enabling standardised comparison through Cohen's Kappa. We employed Bayesian hierarchical modelling to test non-inferiority of medic-LLM agreement compared to medic-medic agreement, using Beta distributed likelihood functions with weakly informative priors. Non-inferiority margins of 0.05, 0.10, and 0.15 were assessed with 95% posterior probability thresholds. Results: Cohen's Kappa for inter-rater reliability was 0.752 (95% CI: 0.710 - 0.794) between medical professionals and 0.391 (95% CI: 0.362-0.420) between LLMs and medical professionals. Bayesian analysis showed medic-medic agreement (posterior mean 0.813, 95% CI: 0.755-0.860) exceeded medic-LLM agreement (0.659, 95% CI: 0.633-0.684) by 0.154 (95% CI: 0.091-0.209). Non-inferiority was rejected at all tested margins (delta = 0.05, 0.10, 0.15). Agreement varied by clinical category, with smallest differences for test results and largest for diagnoses. GPT-4o achieved 97.0% precision and 84.9% recall, with a 7.5% false positive rate. Conclusions: Current LLMs do not achieve inter-rater reliability levels comparable to medical professionals in clinical information extraction from ENT documentation. These findings provide evidence-based guidance for LLM deployment in clinical documentation workflows, suggesting they are best suited for initial extraction with human verification rather than autonomous operation.
L. Barrett, N. Joshi, A. S. North et al.· medRxiv· 0 citations
Diagnostic errors, including misdiagnoses and delayed clinical diagnoses, could affect outcomes of a significant patient population, particularly individuals presenting with rare diseases or non-specific symptoms. From rule-based diagnostic decision supporting systems (DDSS) to large language model (LLM) based tools for clinical reasoning have been developed to address these limitations. However, existing DDSS are often proprietary and difficult to integrate, and recent LLM-based tools remain hindered by operational challenges such as cost, resources constraint, and privacy concerns. Moreover, existing systems interpret electronic medical records (EMR) and generate diagnoses separately, limiting continuous evidence-based analysis and imposing repeated clinician involvement. In this paper, we present DDx-Finder, an open-source framework that leverages Model Context Protocol (MCP) servers for direct EMR and literature access, enabling prompt-driven clinical state extraction and reliable case-report re- trieval via generating searching query by LLM, while addressing limitations related to resource demands and privacy concerns. A clinical case study demonstrates the systems feasibility and its potential to provide accessible, transparent, and systematic differential diagnostic support for complex cases.
H. Lim, H. Yi, J. Y. Yoon et al.· medRxiv· 0 citations
Timely and highly accurate diagnoses by physicians play a crucial role in improving the quality and effectiveness of patient treatment outcomes. Currently, the use of artificial intelligence capabilities in this area has garnered the attention of many health science researchers. Therefore, the main goal of this preliminary study was to compare the text-based diagnostic reasoning performance of emergency medicine physicians and large language models in definitive and differential diagnoses using standardized clinical vignettes. This descriptive comparative study evaluated the diagnostic accuracy of 10 emergency medicine physicians and 4 large language models (LLMs)—ChatGPT (GPT-5.2), Gemini 3, Microsoft Copilot (GPT-4), and Claude Opus 4.1—using 10 standardized clinical vignettes. All LLMs were accessed via official web interfaces. Clinical vignettes were developed from emergency department presentations, validated by an expert panel, and presented as text-only inputs to all evaluators. Diagnostic accuracy was assessed using a standardized scoring protocol: definitive diagnoses required an exact match with expert-derived reference standards, while differential diagnoses required ≥ 3 matches. Overall diagnostic accuracy was the primary outcome. Data were analyzed using Pearson’s Chi-square test, McNemar’s test, and Generalized Estimating Equation (GEE) logistic regression with Bonferroni correction (SPSS version 28). A total of 280 diagnostic evaluations (10 clinical cases assessed by 14 evaluators across 2 diagnosis types) were analyzed. Overall diagnostic accuracy was 61.79%. AI models demonstrated significantly higher overall accuracy (73.75%) compared to emergency medicine physicians (57.00%, p = 0.014). Across all evaluators, definitive diagnoses were more accurate than differential diagnoses (70.71% vs. 52.86%). Generalized Estimating Equation (GEE) analysis revealed a significant interaction between evaluator group and diagnosis type (p = 0.036). Specifically, physicians experienced a significant decline in accuracy when providing differential diagnoses compared to definitive diagnoses (45.0% vs. 69.0%, p = 0.003; remained significant after Bonferroni correction). In contrast, AI models maintained consistently high accuracy across both diagnosis types, with no significant difference between definitive (75.0%) and differential (72.5%) diagnoses (p = 1.000). Large language models outperformed emergency medicine physicians in overall diagnostic accuracy and demonstrated superior consistency across different types of diagnostic tasks. While human physicians struggled significantly with differential diagnoses, AI models maintained high and stable performance regardless of the diagnostic complexity. These findings indicate that AI possesses robust pattern-recognition and reasoning capabilities, suggesting it could serve as a highly reliable clinical decision support tool, particularly in complex scenarios requiring differential diagnostic reasoning. Due to study limitations, such as the small number of clinical scenarios and assessors, these findings should be interpreted with significant caution.
Mehdi Arzani Shamsabadi, Roya Vatankhah, Hasan Jalilvand et al.· International Journal of Eme...· 0 citations
A single LLM verifier lacks sufficient reliability to serve as a stand-alone judge of clinical reasoning at scale, and that structured human oversight remains essential.
Hyunjung Byun, Dahyoun Lee, Munyoung Jung et al.· Journal of medical systems· 1 citation· ⚡1