Skip to content
Book Open access

Does a general-purpose large language model improve physicians’ clinical reasoning?

TL;DR

It is found that LLM access enhances performance on standardized clinical vignettes in all three countries, and policymakers should prioritize structured integration of LLMs as decision-support tools, combined with targeted training, local validation, and safeguards against automation bias rather than relying on access alone.

Abstract

This paper investigates whether access to a general-purpose large language model (LLM) improves physicians’ clinical reasoning across diverse healthcare contexts, as well as the possible implications of using an LLM in healthcare settings. Using a randomized controlled trial with 249 physicians in Indonesia, Kenya, and the Netherlands, the study finds that LLM access enhances performance on standardized clinical vignettes in all three countries. The magnitude of improvement varies, with the largest gains observed in Kenya (+18%), followed by Indonesia (+10.7%) and the Netherlands (+7.2%).The results, however, reveal substantial heterogeneity. Performance distributions overlap, and some physicians with LLM access perform worse than those without, indicating that access alone does not guarantee improvement. Higher usage is associated with better outcomes, and less specialized physicians appear to benefit more, implying that LLMs may help reduce skill gaps. Importantly, the findings emphasize that LLMs function as complements rather than substitutes for clinical expertise. However, the study identifies important risks, including automation bias, hallucinations, and context misalignment, underscoring the importance of careful integration, training, and governance.The paper concludes that while LLMs can enhance clinical reasoning, their effectiveness depends critically on how they are implemented within healthcare systems. The paper recommends that policymakers prioritize structured integration of LLMs as decision-support tools, combined with targeted training, local validation, and safeguards against automation bias rather than relying on access alone. It also emphasizes the need for investment in infrastructure, continuous monitoring, clear liability frameworks, and inclusive governance to ensure equitable, safe, and context-appropriate deployment. It argues for the importance of social dialogue in managing the process.

Read PDF

Similar papers

Review Aug 2026

How large language models can be used for teamwork and communication in healthcare settings: A scoping review.

LLMs hold substantial potential to enhance healthcare teamwork by supporting clinical decisions, streamlining administrative workflows, and improving patient communication, however, ethical, legal, and accountability concerns remain.

Ilse Super, Olya Rezaeian, Onur Asan · 0 citations
Open access Jul 2026

NigBench: A multilingual point-of-care medical query benchmarking study of large language models in Nigeria

A novel benchmark comprising over 9,000 real-world, point-of-care, multilingual, and multimodal clinical question-answer pairs sourced from frontline health workers in Nigeria reveals several critical insights into the suitability of LLMs as clinical decision support systems in low-resource contexts.

Tobi Olatunji, C. Aka, C. Okocha et al. · 0 citations
Conference Jul 2026

A Multi-Domain Human Expert Evaluation of Clinical and Behavioral Knowledge in Large Language Models

Large Language Models (LLMs) such as ChatGPT and Gemini are increasingly used to answer medical and psychological questions, yet systematic evaluations across domains with differing reasoning demands remain limited. We present a multi-domain expert evaluation of two state-of-the-art models, ChatGPT Pro (v5.2) and Gemini 3 Pro, across three healthcare domains: Gynecology, Pathology, and Psychology. We curated 300 open-ended, realistic questions, 100 per domain, designed to elicit clinical reasoning, mechanistic interpretation, and conceptual explanation. Responses were independently scored by domain experts using a standardized five-point rubric. Results reveal domain-dependent performance patterns, with both models performing well in guideline-aligned advisory tasks in Gynecology, lower in mechanistic diagnostic contexts in Pathology, and showing divergent strengths in conceptual psychology questions. Quantitative and qualitative analyses highlight recurring limitations in contextual nuance, mechanistic depth, and safety framing. These findings underscore the importance of domainstratified evaluation and expert oversight when deploying LLMs for healthcare information and provide a reproducible framework for future assessments.

Pragna Prahallad, Pranathi Prahallad, Dhrithi P. Desai · 0 citations
Review Open access Aug 2026

Large Language Models for Differential Diagnosis: A Survey of Performance, Collaboration, and Technical Strategies

Errors in differential diagnosis often arise while clinicians are generating and comparing candidate explanations. This review examines the use of large language models (LLMs) for this part of diagnostic reasoning. Internal medicine and pediatrics are the main focus; evidence from radiology, surgical subspecialties, infectious disease, and mental health is used to examine how findings change across specialties. Reported performance depends on the clinical setting, the quality of the input, the prompt, model adaptation, and the evaluation design. Some studies place LLMs near trainees and find that they produce wider, better-organized differentials. Experienced clinicians, however, remain more reliable overall. Domain adaptation, external knowledge, and interactive workflows have improved performance in specific evaluations, but hallucinations and automation bias remain, alongside unresolved questions of governance. Current evidence therefore supports clinician-supervised use of artificial intelligence (AI) systems rather than autonomous diagnosis, pending prospective and specialty-specific evaluation.

Yun-Jia Wu, Qi Yan, Dingcheng Tian · 0 citations
Review Open access Jul 2026

Medical students' use of large language models: a national survey.

BACKGROUND Large language models (LLMs) are increasingly embedded in medical education and clinical care settings, yet contemporary Canadian data describing medical students' use and perceptions remain limited. OBJECTIVE To quantify the prevalence, frequency, and patterns of LLM use among medical students in Canada; to characterize perceptions of utility, accuracy, limitations, and impact; and to describe perceived barriers, challenges, and ethical/privacy concerns. METHODS We conducted a national, cross-sectional survey distributed to English-speaking medical students between November and December 2025. Recruitment occurred through medical school channels, student unions, and national/regional student organizations. RESULTS Among 286 respondents from 10 medical schools, 96.50% reported using at least one LLM. The most commonly used LLMs were ChatGPT (93.36%) and OpenEvidence (57.69%). Daily/weekly use was most frequent for coursework assistance (60.22%) and clinical questions (57.14%). Most respondents reported positive impacts on efficiency (81.62%), learning (77.01%), and academic performance (59.49%). Students commonly reported encountering inaccurate information (90.18%). Formal instruction on LLM use was uncommon (10.95%), though 67.67% of students agreed medical schools should integrate formal instruction on LLMs. Only 21.43% of respondents felt adequately educated on data privacy regulations applicable to these tools. CONCLUSION While LLM use among surveyed medical students in Canada was nearly universal and perceived favorably, students reported exposure to inaccurate outputs and substantial gaps in formal training and privacy literacy. These findings support the development of structured curricular guidance on appropriate application of these tools, including information verification practices and ethical, privacy-aware engagement.

Austin A. Barr, R. Rozman, Kevin Liu et al. · 0 citations
Review Jul 2026

Large language models in clinical and healthcare scenarios: a global informatics analysis

This paper conducts a comprehensive analysis of evaluation methods, deployment processes, and governance strategies for LLMs in the healthcare field, focusing on three key issues: model version drift, multilingual external validation, and prompt injection security governance.

Song-Bin Guo, Sui-Xing Zhong, Yixian Ma et al. · 0 citations