Skip to content
Review Open access

Interactive platform for supporting clinical decision-making using large language models (LLMs)

Jun 2026 · INNOVATIVE TECHNOLOGIES AND SCIENTIFIC SOLUTIONS FOR INDUSTRIES · pp. 173-184 · 0 citations

TL;DR

The results suggest that specialized, locally deployable fine-tuned models, combined with bounded interaction design, standardized output constraints, and workflow-oriented system integration, provide a secure, practical, and effective pathway for incorporating LLM-based decision support into routine pre-ambulatory clinical workflows while preserving safety, usability, and auditability in practice.

Abstract

Current healthcare systems face increasing workload, fragmented communication, and documentation burden, which contributes to delays and diagnostic errors in early triage. The proposed solution is intended to improve consistency, speed, and standardization in early patient assessments overall. This study presents an interactive clinical decision support platform that operationalizes a workflow-constrained, two-step history-taking process to support symptom-based differential diagnosis in the pre-ambulatory phase of care. The system addresses three tasks: (T1) automated patient history elicitation via constrained dialogue (exactly two multiple-choice follow-up questions), (T2) formalization of symptom narratives into structured medical (Latinate) terminology for clinician-facing documentation, and (T3) generation of differential diagnosis recommendations under a fixed and clinically interpretable output schema. We compare a generic large language model baseline (GPT-4, zero-shot) with a domain-adapted model (Llama-3 fine-tuned using LoRA) under identical interaction, prompting, and formatting constraints. Experiments on 903 symptom–diagnosis records and a held-out set of 200 controlled vignettes show that domain adaptation yields a 10–12% macro-F1 improvement and approximately 15% higher Recall on complex cases, while producing more clinically discriminative and diagnostically relevant follow-up questions as assessed by two medical raters . In a workflow-level evaluation, the platform reduced documentation time by 28% compared to standard intake and lowered the administrative effort required to prepare an initial clinician-facing summary for physician review. These results suggest that specialized, locally deployable fine-tuned models, combined with bounded interaction design, standardized output constraints, and workflow-oriented system integration, provide a secure, practical, and effective pathway for incorporating LLM-based decision support into routine pre-ambulatory clinical workflows while preserving safety, usability, and auditability in practice.

Read PDF

Similar papers

Open access Aug 2026

A Human-in-the-Loop Large Language Model System Based on the Model Context Protocol for Differential Diagnosis from Electronic Medical Records and Literature

Diagnostic errors, including misdiagnoses and delayed clinical diagnoses, could affect outcomes of a significant patient population, particularly individuals presenting with rare diseases or non-specific symptoms. From rule-based diagnostic decision supporting systems (DDSS) to large language model (LLM) based tools for clinical reasoning have been developed to address these limitations. However, existing DDSS are often proprietary and difficult to integrate, and recent LLM-based tools remain hindered by operational challenges such as cost, resources constraint, and privacy concerns. Moreover, existing systems interpret electronic medical records (EMR) and generate diagnoses separately, limiting continuous evidence-based analysis and imposing repeated clinician involvement. In this paper, we present DDx-Finder, an open-source framework that leverages Model Context Protocol (MCP) servers for direct EMR and literature access, enabling prompt-driven clinical state extraction and reliable case-report re- trieval via generating searching query by LLM, while addressing limitations related to resource demands and privacy concerns. A clinical case study demonstrates the systems feasibility and its potential to provide accessible, transparent, and systematic differential diagnostic support for complex cases.

H. Lim, H. Yi, J. Y. Yoon et al. · 0 citations
Open access Jul 2026

Augmenting medical data interpretation with Large Language Models (LLMs): a comparative analysis of patient empowerment, information processing, and technology acceptance.

BACKGROUND Medical data interpretation traditionally relies on healthcare professionals as intermediaries, which can limit patient autonomy and engagement. Large Language Models (LLMs) present an opportunity to transform this paradigm by enabling direct patient access to AI-generated interpretations; however, comparative research on their effectiveness across different medical data types and communication modalities remains limited. This study explores how direct LLM-augmented interpretation of medical data, in which patients use an AI system to receive real-time explanations of laboratory and radiological results, compares with healthcare professional-led interpretation across different data modalities, with particular attention to patient comprehension, empowerment, and technology acceptance. METHODS Using a mixed-methods approach with a within-subjects experimental design, 45 demographically diverse participants experienced six scenarios: blood work and medical imaging interpretations delivered via (1) healthcare professional phone consultation, (2) in-person consultation, or (3) LLM interaction through a custom-configured ChatGPT-4o interface (Medical Explainer AI) designed to provide plain-language explanations of findings, highlight abnormal values, contextualize clinical significance, explain medical terminology, and adapt explanation complexity based on user feedback. RESULTS LLM interaction significantly enhanced diagnostic comprehension (mean difference = 1.3 compared to phone consultation, p < 0.001), reduced cognitive load, increased perceived control, and improved time efficiency. Healthcare professional-led interpretation, particularly in-person, maintained advantages in fostering trust, reducing anxiety, and enhancing confidence in decision-making. The benefits of LLM interaction were more pronounced for blood work than for medical imaging interpretation. Age, education level, and health literacy significantly moderated the effectiveness of different interpretation methods. CONCLUSIONS LLMs offer complementary rather than replacement capabilities for medical data interpretation, excelling in enhancing comprehension, control, and efficiency, while healthcare professionals provide superior relational value through trust, confidence, and emotional support. Implementation strategies should leverage the strengths of both approaches, carefully considering data complexity and patient characteristics to maximize benefits while ensuring equitable access.

Pouyan Esmaeilzadeh · 0 citations
Review Open access Jul 2026

A Modular Evaluation of AI-Assisted Clinical Documentation

Clinical documentation in Electronic Health Records (EHRs) remains a substantial source of administrative burden for clinicians. In this study, we evaluate a modular AI-assisted clinical documentation pipeline using two complementary approaches: (1) a controlled benchmark based on multilingual synthetic clinical dialogues, and (2) an observational analysis of real-world usage traces from routine deployments. The benchmark enables systematic comparison of ASR–LLM configurations under fully controlled conditions, using metrics for transcription accuracy (Word Error Rate and Medical WER), report-generation quality, and modeled processing cost. Within this benchmark setting, Voxtral showed the strongest ASR performance among the evaluated models, while GPT-4o and Gemini 1.5 Pro showed the strongest report-generation performance under the automated evaluation used in this study. The real-world trace analysis should be interpreted as descriptive evidence of operational use, not as prospective clinical validation or as a direct evaluation of any single benchmarked configuration. Taken together, the results support the use of this pipeline as a human-supervised draft-generation tool that still requires clinician review, local workflow evaluation, and prospective clinical validation before broader deployment.

Julien Delaunay, Maissaa Sarkis, Jordi Solé-Casals et al. · 0 citations
Review Open access Jul 2026

Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning

A dual-view approach that connects clinical practice with computational methods is presented, establishing a five-level competency scheme following Miller’s Pyramid and linking deductive, inductive, and abductive reasoning patterns to common medical goals and tasks.

Qi Peng, Jiatong Li, Sirui Huang et al. · 4 citations
Review Open access Aug 2026

Large Language Models for Differential Diagnosis: A Survey of Performance, Collaboration, and Technical Strategies

Errors in differential diagnosis often arise while clinicians are generating and comparing candidate explanations. This review examines the use of large language models (LLMs) for this part of diagnostic reasoning. Internal medicine and pediatrics are the main focus; evidence from radiology, surgical subspecialties, infectious disease, and mental health is used to examine how findings change across specialties. Reported performance depends on the clinical setting, the quality of the input, the prompt, model adaptation, and the evaluation design. Some studies place LLMs near trainees and find that they produce wider, better-organized differentials. Experienced clinicians, however, remain more reliable overall. Domain adaptation, external knowledge, and interactive workflows have improved performance in specific evaluations, but hallucinations and automation bias remain, alongside unresolved questions of governance. Current evidence therefore supports clinician-supervised use of artificial intelligence (AI) systems rather than autonomous diagnosis, pending prospective and specialty-specific evaluation.

Yun-Jia Wu, Qi Yan, Dingcheng Tian · 0 citations
Review Jul 2026

Large language models in clinical and healthcare scenarios: a global informatics analysis

This paper conducts a comprehensive analysis of evaluation methods, deployment processes, and governance strategies for LLMs in the healthcare field, focusing on three key issues: model version drift, multilingual external validation, and prompt injection security governance.

Song-Bin Guo, Sui-Xing Zhong, Yixian Ma et al. · 0 citations