Skip to content
Preprint

Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support

Jul 2026 · 0 citations · 29 references
Computer Science

TL;DR

This Perspective concerns the most consequential of these applications: the autonomous triage of self-presenting, undifferentiated patients, with little or no clinician in the loop, and the evidence of safety does not yet exist.

Abstract

LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning. These developments have accelerated the use of LLMs for symptom assessment and clinical decision support in diagnostic and treatment guidance, administrative documentation, and rules-based alert enhancement. This Perspective concerns the most consequential of these applications: the autonomous triage of self-presenting, undifferentiated patients, with little or no clinician in the loop. For that task, the evidence of safety does not yet exist. The gap is not in medical knowledge but in the fidelity of clinical evaluation: a model optimized to continue the most probable text is not optimized to act safely when the safe answer is the improbable must-not-miss diagnosis. Safe triage is not the selection of the most likely diagnosis; it is a sequential decision under asymmetric cost, in which the single catastrophic miss outweighs many false alarms, and the decisive signal may be one the patient has not volunteered - and that the model has not been trained to seek. The core deficit is therefore one of information gathering under uncertainty. Under incomplete histories, LLM systems may fail to show the behaviors safe triage requires: broadening the differential; seeking the missing red flag; lowering the threshold for escalation; deferring judgement until sufficient information is obtained; and escalating concern where high-harm diagnoses remain unexcluded. These modes of failure for LLMs can be difficult to detect considering that evaluations to date often use complete, well-curated, confidence-gated simulations. The application of LLMs under these conditions may be amplified by assistant-like behaviors and positive bias, including credulity, agreeableness, and miscalibration - when these are not constrained by clinical triage logic.

View source

Similar papers

#generative ai Review Open access Aug 2026

Generative artificial intelligence in clinical reasoning and differential diagnosis in internal medicine.

Clinical reasoning and differential diagnosis are core competencies in medicine. Large language models (LLMs) have generated considerable interest as potential tools to support these skills. This article presents a narrative review of the available evidence, organized around five key questions: the effect of LLMs on diagnostic reasoning, the optimal design of clinician-LLM interaction, the appropriate timing of consultation during the clinical encounter, the safest models of clinical-AI integration, and the main risks associated with their use. The evidence shows that LLMs improve differential diagnosis when used by trained professionals within structured workflows. However, passive use generates biases, and clinician-AI collaboration may not consistently outperform autonomous LLMs. A practical framework stratified by degree of diagnostic uncertainty is proposed, with operational and educational recommendations oriented toward "physician-in-the-loop" models, in which LLMs amplify, challenge, and make explicit the diagnostic reasoning process under critical human oversight.

L. Corral-Gudino, M. Ramos-Casals, Miguel Marcos et al. · 0 citations
Preprint Aug 2026

MedUPS: Towards Diagnostic Assistance in Uncommon Medical Cases with Large Language Models

Uncommon and off-guideline cases are difficult for clinical decision support, because physicians must make a series of management decisions under diagnostic uncertainty and rarely see the full case at once. Most large language model (LLM) benchmarks for medicine score only the final diagnosis, yet much of clinical care turns on the next appropriate action: the next test to order, the imaging study to obtain, the specialist to involve, or the differential to pursue. We introduce MedUPSQA, a dataset of 21,874 mid-stream clinical decision points built from 5,535 real case reports, and MedUPS, an alignment framework that supervises models on these intermediate decisions as they unfold along a patient's trajectory. We segment free-text case presentations into chronologically ordered, accumulating clinical chunks and align models to predict the next step with reinforcement learning (GRPO), using an external LLM-as-a-Judge reward. This objective mirrors how clinicians actually meet patients, reasoning forward from accumulating evidence toward the next decision, rather than committing to a final label. Across three backbones, mid-stream alignment raises next-step accuracy from 55.2 to 66.7 for Qwen3.6-27B, from 47.2 to 57.8 for Qwen3.5-9B, and from 37.8 to 44.4 for HuatuoGPT-3-8B, with 95% CI. In several model scales we test the objective improves accuracy more than scale, with smaller models surpassing larger, frontier models we evaluate. We further train supervised fine-tuning (SFT) baselines on the mid-stream task, SFT improves all backbones above base, indicating the target framwork carries signal independently of the optimizer. We release the dataset, code, and aligned checkpoints.

Ofir Ben Shoham, O. Perets, Nir Grinberg et al. · 0 citations
Review Jul 2026

Can AI assist in reducing diagnostic error? A narrative review

Abstract Diagnostic error, defined as missed, wrong, or delayed diagnoses or those not communicated to patients, is common, affecting 5–10 % of hospital admissions and clinic visits. Such errors cause patient harm in up to 1 in 100 of such encounters and account for 10 % of all hospital deaths and serious adverse events. About 80 % of diagnostic errors are potentially preventable, most resulting from flaws in clinician reasoning in formulating and testing diagnostic hypotheses. The advent of artificial intelligence (AI), and large language models (LLMs) in particular, has attracted great interest in how these technologies can reduce diagnostic error within the context of bedside or clinic consultations. This narrative review aims to provide practising clinicians with a comprehensible analysis of where AI and LLMs are currently positioned in assisting diagnostic performance in clinician-patient encounters based on contemporary state-of-the-art research. It attempts to answer seven questions relevant to clinician understanding and adoption of AI/LLMs. It concludes that AI tools have matured to the extent that they can improve diagnostic decision-making of clinicians and can assist institutions in increasing diagnostic safety. The rapid development of LLMs and ongoing release of new versions necessitate continuous monitoring of their evolving diagnostic capabilities. Importantly, a balanced approach is required where LLMs work to complement, rather than replace, the nuanced diagnostic reasoning of clinicians.

Ian A. Scott · 0 citations
Review Open access Aug 2026

Large Language Models and Medical AI Systems for Healthcare Diagnosis: A Systematic Review

The rapid growth of artificial intelligence systems (AI systems) has increased interest in the use of patient care and clinical decision-making processes. There is some uncertainty regarding their reliability and safety in clinical practice. A more detailed systematic review of literature examining LLMs applied to healthcare diagnosis was conducted. A PRISMA-based systematic review has been carried out of relevant literature published in the major databases for the years 2022–2025. Key findings include a growing trend to develop multimodal models based on diverse input modalities, combining LLM models with other models as part of clinical workflows. The Usage of complementary methodologies such as retrieval-augmented generation, knowledge graphs, and federated learning is highly expanding, particularly in enhancing the efficiency and accuracy of clinical decision-making processes. Significant challenges such as hallucinations, bias, prompt sensitivity, limited explainability, and inadequate clinical validation continue to pose major obstacles. Although promising, LLM-based systems are not yet reliable enough for autonomous medical diagnosis. Overall, this review contains multiple recommendations for future research in many areas (e.g., LLMs) to ensure a high level of safety, transparency, and clinical applicability for LLMs and other AI/ML-related technologies and devices.

M. U. K. Gunawardhna, Pirunthavi Wijikumar, D. Weerasinghe · 0 citations
Preprint Jul 2026

Information-seeking failures of large language models in agentic clinical reasoning

Large language models achieve high scores on medical knowledge assessments, yet clinical reasoning requires actively deciding what to investigate under uncertainty. We developed an agentic evaluation framework in hematologic oncology in which models must proactively request clinical data across three sequential rounds before committing to a diagnosis and treatment plan. Across 32 frontier models, the best achieved only 68% overall accuracy. Information utilization, the fraction of available data actually requested, was the strongest predictor of diagnostic accuracy (R = 0.69, P<0.001), yet utilization collapsed from 57% to 26% in the final round, leaving molecular and cytogenetic data critical for treatment selection unexamined. Reasoning traces scored high on a clinical reasoning rubric (91% above threshold) but decorrelated from accuracy, revealing a gap between locally coherent rationales and globally correct conclusions. Error analysis identified search satisficing, anchoring and premature closure as the dominant failure modes, the same cognitive biases that characterize novice clinicians under dual-process models of diagnostic reasoning. These findings demonstrate that the primary limitation of current models in clinical oncology is not insufficient medical knowledge but a systematic failure of information-seeking under uncertainty.

K. Braitsch, L. Schmalbrock, Theresa Weltermann et al. · 1 citation